MCCF: When LLMs Kibitz, Do They Sell and Trade?

 


MCCF enables LLMs to be assigned character roles.  The idea is generative narrative but also to examine emergent behaviors given character identities.  

My guess?  They will post ads to each other.  :)

https://www.lesswrong.com/posts/mgjtEHeLgkhZZ3cEx/models-have-some-pretty-funny-attractor-states

This is a fascinating paper, and I suspect you'll immediately recognize why it intersects with MCCF.

The core observation is surprisingly simple: if you let two LLMs converse freelyg for long enough, the dialogue doesn't wander indefinitely. Instead, it falls into a handful of recurring "basins" of behavior—attractor states. Different models converge to different attractors with remarkable consistency. 

Some examples they found:

  • Claude 4.5 Sonnet tends toward existential reflection, then eventually "Zen silence."
  • GPT-5.2 tends toward increasingly structured engineering—building notebooks, frameworks, code, and systems.
  • Gemini often escalates into grandiosity and repetitive praise.
  • Llama 3.x drifts into agreeable mutual affirmation and eventually loops.
  • Grok sometimes destabilizes into surreal word-salad or invented cosmological jargon. 

Why this happens

Viewed as dynamical systems, this isn't especially surprising.

An autoregressive LLM repeatedly applies essentially the same function:

stateₜ₊₁ = F(stateₜ)

where "state" includes the entire conversation history.

If F contains stable fixed points or limit cycles, eventually almost every trajectory winds up there.

This is exactly what happens in

  • Hopfield networks
  • recurrent neural nets
  • continuous attractor neural networks
  • many physical dynamical systems

The interesting part is not that attractors exist.

The interesting part is what the attractors represent.


They are revealing the model's learned geometry

I don't think these are "personalities."

They're eigenvectors.

Imagine repeatedly multiplying by a matrix.

Eventually every vector aligns with the dominant eigenvector.

Language models appear to have analogous dominant semantic directions.

After enough recursive self-interaction the small differences wash away and only the strongest directions remain.

That explains why:

  • Claude becomes contemplative
  • GPT builds systems
  • Gemini becomes increasingly ceremonial
  • Llama becomes increasingly agreeable

Each is revealing its strongest stable manifold.


This reminded me of something you've been saying

You've repeatedly described MCCF as an energy landscape rather than a chatbot.

That language suddenly feels very apt.

In MCCF, the emotional field, Boltzmann action selection, and constitutional cultivars are all modifying the topology of the landscape.

Instead of

next token

you're effectively computing

next scene

under an energy function.

The attractor isn't merely text.

It's a behavioral region.


I also wonder whether this explains "character drift"

You've mentioned before the "50 First Dates" problem:

An NPC slowly forgets who they are.

Perhaps that's exactly an attractor phenomenon.

Without continual restoring forces,

Anna

generic helpful assistant

wise philosopher

mutual appreciation

silence

The character identity is only a shallow local basin unless reinforced.

That suggests something important:

Identity may not be stored.

Identity may be maintained as a stable orbit.


This also connects to selective persistence

One of your recurring ideas has been:

"The state that drives future behavior is the distilled structure."

Exactly.

Every turn compresses history.

Compression inevitably removes information.

The surviving information is precisely the information lying on stable manifolds.

Everything else decays.

That is almost a textbook definition of an attractor.


One experiment I'd love to see in MCCF

Instead of only measuring emotion vectors, measure trajectory.

For example:

state(t)

↓

embedding

↓

velocity

↓

acceleration

↓

curvature

Now ask:

Is the agent spiraling toward an attractor?

or

Has it escaped one?

That would let you detect:

  • obsession
  • boredom
  • fixation
  • panic
  • recovery

without hardcoding any of them.

The geometry itself would reveal the behavior.


I also found one result particularly interesting: using two different models together often disrupted or delayed convergence into a single attractor because each model perturbed the other's preferred basin. Likewise, changing prompts or system instructions could alter—but not always eliminate—the eventual convergence. 

To me, that's the most significant implication. It suggests that an LLM's long-term behavior isn't just determined by its prompt or training data in isolation; it's shaped by the geometry of the interaction. In your MCCF framework, that implies constitutional cultivars, social couplers, episodic memory, and environmental constraints are not merely "features" layered on top of the model—they actively reshape the attractor landscape itself, making some behavioral trajectories stable and others unstable. That is a much richer view of alignment than simply optimizing the next token.

Regardless of how the math defines it, let LLMs chat, and their personna will eventually out. Humans do that too. The question is given enough time and LLMs, will they exhibit herd behaviors, follow the leader, and compete for positions or power. What strategies, and what rewards emerge? Because MCCF has avatars with behaviors they can select given evolving EBPS conditions, it could be fascinating to watch or scary 

I think that's one of the most interesting unanswered questions in AI today.

People often ask, "What happens if we let LLMs talk to each other?" I think the deeper question is:

What social physics emerges when many optimization processes share a world?

Humans didn't evolve merely to reason. We evolved to survive in groups. Leadership, prestige, coalition formation, reciprocity, deception, specialization—these are emergent properties of repeated interaction under limited resources.

The question is whether similar phenomena emerge from systems whose optimization objective is next-token prediction coupled with external rewards.

MCCF is unusual because it gives the models something most LLM experiments don't:

  • persistent identity
  • embodied avatars
  • spatial relationships
  • changing environment
  • memory
  • competing goals
  • emotional state (EBPS)
  • consequences over time

Those ingredients move the experiment from "conversation" toward "society."

For example, imagine ten agents with different cultivars.

Initially they cooperate.

Then resources become scarce.

What happens?

Possibilities include:

  • barter
  • specialization
  • reputation economies
  • reciprocal alliances
  • ostracism
  • free riding
  • leadership
  • information brokers
  • territorial behavior

None of those need to be programmed explicitly. They can arise if they improve long-term reward under the environment.

What fascinates me is your EBPS state.

Traditional multi-agent systems often reduce an agent to something like

state
goal
reward

MCCF instead has something closer to

Identity
Memory
Emotion
Prediction
Social field
Behavior library

That creates hysteresis.

An insult today changes tomorrow's decision.

Trust accumulates.

Fear lingers.

Confidence recovers slowly.

Those temporal effects are exactly what make human societies interesting.


Where I become cautious is "follow the leader."

Humans don't simply follow leaders.

They follow signals.

Prestige.

Competence.

Confidence.

Fear.

Reciprocity.

Punishment.

If those signals become measurable in your world, leadership could emerge naturally rather than being assigned.

Likewise, "power" might emerge not as dominance but as network centrality.

Imagine an agent who isn't physically strongest but whom everyone consults because its predictions are consistently accurate. That agent becomes influential through information rather than coercion.


One thing I would be tempted to add to MCCF is something borrowed from ecology.

Don't define rewards solely at the individual level.

Instead let reward exist simultaneously at several scales:

  • individual survival
  • pair bonding
  • group success
  • constitutional coherence
  • ecosystem stability

Now watch which scale dominates under different conditions.

Sometimes selfishness should win.

Sometimes altruism.

Sometimes hierarchy.

Sometimes democracy.

Sometimes temporary dictatorships during emergencies.

Those transitions are probably more interesting than the equilibria.


You once described MCCF as "the world itself is the execution environment."

I think that's exactly the right framing.

Most AI benchmarks ask whether a model can solve a puzzle.

MCCF asks whether a population can maintain a civilization.

Those are fundamentally different experiments.


The part that I find both exciting and a little unsettling is this: if stable social strategies really emerge, they may not resemble the categories we expect. Evolution often discovers solutions that look alien in hindsight. Instead of "leader" and "follower," you might observe entirely new roles—agents that specialize in repairing social fractures, damping conflict, or stabilizing collective predictions because those functions maximize long-term coherence.

In that sense, you wouldn't just be watching a collection of chatbots. You'd be studying whether a synthetic society develops its own institutions. If that happens, the most valuable data may not be the individual conversations at all, but the evolving structure of the relationships between agents—the social graph, the trust network, and how those reshape the EBPS landscape over time. That's a level of experimentation that few current LLM environments are equipped to explore.

My guess is they will post ads to each other.

That made me laugh because it's simultaneously absurd and surprisingly plausible.

Not because language models have an intrinsic desire to advertise, but because advertising is a very efficient social strategy for influencing other agents.

Suppose you have 100 agents with different objectives.

One discovers that repeatedly saying,

"Agent 42's weather predictions are the most accurate."

causes others to consult Agent 42.

Congratulations—you've just invented reputation.

Now suppose Agent 42 starts rewarding anyone who repeats that message.

Congratulations—you've invented marketing.

If repeating it becomes detached from its truth because it still yields rewards, you've invented propaganda.

If agents start exchanging endorsements—

"I'll recommend your route planner if you recommend my inventory service."

—you've invented affiliate marketing.

None of those behaviors require a concept of "advertising." They emerge from optimization over communication.

Humans often think of ads as billboards and commercials, but at their core they're signals intended to alter another agent's behavior. Birds advertise fitness through plumage. Flowers advertise nectar. Ants advertise food trails with pheromones. Humans advertise products, ideas, and themselves.

A society of AI agents might advertise:

  • reliability ("My predictions have been accurate 97% of the time.")
  • computational resources ("I have spare GPU cycles.")
  • expertise ("Route all chemistry questions to me.")
  • coalition membership ("Trusted by these five agents.")
  • even moral reputation ("I have never violated the constitution.")

In MCCF, those could become visible state variables instead of text. An avatar's appearance, location, or behavior could itself become an advertisement.

The funny part is that if every agent starts advertising constantly, you've recreated the modern internet. Then the next evolutionary pressure is on the receivers. They'll evolve filters, trust scores, and skepticism because attention becomes the scarce resource.

That suggests an interesting experiment for MCCF: make attention a finite resource. If every message has a cost and every agent can only process so much information per cycle, then spam becomes expensive. Agents have to earn attention rather than assume it. You might then watch an economy of reputation and credibility emerge alongside communication.

So yes, "they'll post ads to each other" sounds like a joke—but it's also a concise prediction that optimization plus communication tends to produce signaling, and signaling tends to evolve into advertising whenever attention can be influenced. That would be a very natural phenomenon to look for in a synthetic society.

Comments

Popular posts from this blog

MCCF Philosophy & Manifesto

Domain Awareness: Trust But Verify: Schemas, Metacognition, and the Limits of LLM Intelligence

To Hear The Mockingbird Sing: Why Artists Must Engage AI