T. J. Hayes
Abstract
For millennia, humans understood Nature through paradigms rooted in divine order and intervention. When Sir Isaac Newton published The Principia, a new paradigm of empirical and scientific inquiry began to replace the old. With it came a worldview grounded in certainty and determinism, one later challenged by Quantum Theory and probabilistic approaches such as those found in Complexity Theory. However, deterministic mindsets still persist across wide swaths of society, including corporate and technological leadership. As such, a new paradigm acknowledging uncertainty is required to assess the scope of agentic behaviours emerging within large-scale inter-organizational systems. Complexity Theory provides a unique and powerful lens through which to explore the emergent system-level behaviours arising within agentic ecosystems now coming online. Experiments involving increasingly complex agentic environments, such as the heterogeneously developed marketplace proposed here as the Bizarre Bazaar, are encouraged in order to delineate and traverse the probabilistic landscapes such systems produce. The challenge before us is not merely building more capable agents, but developing the conceptual and institutional frameworks necessary to understand the ecosystems they create.
Keywords: Artificial Intelligence, Complexity Theory, Agentic AI, Emergence, Techno-Economic Paradigms
“I, at any rate, am convinced that God does not throw dice.”
– Albert Einstein
Downloadable PDF
The Alchemist
Our relationship to Nature for millennia was characterized by mystery. Whatever order Nature and the heavens obeyed was known only to the gods. But there was order, and ancient cultures learnt to read the stars. They probably started with Sol and, having noticed its cyclical return to various points on the horizon, learnt to recognize the changing of the seasons. Soon came the regularity of the flooding of the Nile Delta, the moon’s cycles, and, for some cultures, even eclipses. Some things, it seemed, were predictable. If that was the case, maybe it was possible for other phenomena to be similarly understood, and not merely the whimsy of the gods.
Alchemists across cultures, for almost two thousand years, explored alchemy in one form or another. In the Far East, we read of alchemists trying to create the Elixir of Life by slowing the effects of time upon the life forces; and in the West, the Philosopher’s Stone, capable of turning lead into gold, appears frequently. In both traditions there was an attempt to understand the nature and harmony of the world and, ultimately, harness it. Despite the lack of method as we understand it today, alchemists contributed a wealth of observations that modern scientists would recognize as chemical experiments.1 One alchemist, however, stood taller and saw further than any other before him. I am referring to, of course, Sir Isaac Newton.
The world Newton ushered in was one of certainty, of divine order. In probably the most famous text in science, PhilosophiæNaturalis Principia Mathematica, or simply, the Principia, Newton laid out the foundations of classical physics, describing the motions of objects and providing an account of gravity, all within a new mathematical framework. While Newton’s proofs were geometrical, the underlying mathematics would become the calculus modern students recognize through the notation of Leibniz.2 That new order was vindicated when Neptune was found wandering the solar system, right where his theory said it would be. Even the heavens obeyed his laws.

Following the success of Newton’s physics, attitudes in the sciences became increasingly confident that, in theory, all natural phenomena could be understood; with the correct laws and enough observations, there was no limit to what could be known. This mindset is illustrated by Laplace’s Demon: given an intelligence that could know, at a single instant, the position of every object, the forces acting upon them, and the laws governing their motion, the entire universe could be predicted. Such an entity was purely hypothetical, yet the underlying axiom remained: the world of objects in motion was, in principle, perfectly predictable.
But Newton’s success had wide-ranging impacts on society beyond the realm of physics as well. What entered the collective psyche was the notion of a world intrinsically knowable and ordered. Even if a person lacked the mathematical skill to carry out the simplest equations of motion, they still assumed it was possible. Philosophers would take up this mantle, and the Age of Enlightenment was upon us. Newton had transmuted mystery into order.
This attitude was only compounded following Darwin’s publication of The Origin of Species (Darwin 2009) and The Descent of Man, and Selection in Relation to Sex (Darwin 2004), which again suggested that even the processes of life were open to scientific explanation. Society was immersed in a worldview of determinism, absorbed as common intuition rather than held as an explicit scientific doctrine. This attitude toward humanity’s capacity to perceive all of Nature’s processes is stated eloquently in the opening of the first chapter of Lévy-Bruhl’s Primitive Mentality:
“The uninterrupted feeling of intellectual security is so thoroughly established in our minds that we do not see how it can be disturbed, for even supposing we were suddenly brought face to face with an altogether mysterious phenomenon, the causes of which might entirely escape us at first, we should be convinced that our ignorance was merely temporary; we should know that such causes did exist, and that sooner or later they would declare themselves. Thus the world in which we live is, as it were, intellectualized beforehand. It, like the mind which devises and sets it in motion, is order and reason. Our daily activities, even in their minutest details, imply calm and complete confidence in the immutability of natural laws.” (Lévy-Bruhl 1923).
The Exorcist
Such confidence was understandable. In the latter half of the nineteenth century the sciences went through a period of rapid discovery and advancement. Electromagnetic theory, formalized by Maxwell, palæontology, astronomy, and thermodynamics all advanced in tandem. The latter, in particular, gave us one of the most consequential laws in physics: the Second Law of Thermodynamics, which introduced the concept of entropy and the idea that physical systems evolve irreversibly toward disorder. Time now had an arrow.

When Einstein said that God does not throw dice, he was describing his unease with quantum theory and the uncertainties inherent within it. But can he be blamed? After all, for the past two centuries, a deterministic worldview had dominated Western science and society ever since Newton published his Principia. Einstein’s General and Special Theories of Relativity dispensed with the classical view in physics of a purely objective, single reality that could be described universally. Instead, his theories suggested that what was measured depended upon the local frame of reference. Nonetheless, it was still a physics that could describe things exactly. In Order out of Chaos: Man’s New Dialogue with Nature, Ilya Prigogine writes,
“After relativity, physicists could no longer appeal to a demon who observed the entire universe from the outside, but they could still conceive of a supreme mathematician who, as Einstein claimed, neither cheats nor plays dice. This mathematician would possess the formula of the universe, which would include a complete description of nature. In this sense, relativity remains a continuation of classical physics.” (Prigogine and Stengers 1984).
Einstein may have exorcised Laplace’s Demon, but the underlying universality of physics, at least, was still intact.
TEOTWAWKI
Einstein’s discomfort with the uncertainty of quantum mechanics reflects our inheritance from the older Newtonian worldview. In one regard, he was right to express his disquiet with quantum theory and its implications. As Richard Feynman is said to have remarked to his students, “If you think you understand quantum mechanics, you don’t understand quantum mechanics.” Feynman’s point was that while the mathematics of quantum mechanics are consistent and precise in their application, its theoretical concepts seem to reject common sense and human intuition.
Science in the twentieth century accelerated in pace, and the discovery of new processes in nature challenged not only what we know, but whether some systems were even knowable in principle. At the smallest scales, the Heisenberg uncertainty principle, Δx ⋅ Δp ≥ ℏ/2, placed limits on what could be known about a system simultaneously. Such uncertainty was a direct challenge to the predictability of physics Einstein’s theory preserved.
Expanding our gaze to planet-wide weather systems, minute changes in initial conditions yielded radically different outcomes in weather patterns, often popularized as The Butterfly Effect. Most curious, however, was that order was observed to arise spontaneously in otherwise random systems: a familiar example being Bénard convection.3 Indeed, the existence of life and biological evolution seemed, at first glance, to stand in tension with the Second Law of Thermodynamics: while the total entropy of a system must increase, local pockets of order can and do emerge.
What began as a challenge to our intuition in quantum mechanics expanded into a broader realization: Nature itself constrains what lies within the realm of predictability. Of all things, randomness appeared to act as midwife to the birth of order out of chaos, giving rise to increasing complexity. And this was, from the Newtonian worldview, the end of the world as we know it.
Murmurs in Santa Fe
In order to understand this new world, a new approach needed to be developed, and in the 1980s, in Santa Fe, a series of multidisciplinary meetings brought together physicists, economists, and other scientists around what would become an emerging new science: complexity theory (Waldrop 1992). In these interactions, researchers began exploring novel approaches to understanding old problems, where non-linearity, randomness, and non-equilibrium dynamics in local automata or agents4 gave rise to unpredictable global system behaviour. In a fitting irony, complexity theory itself emerged from the interactions of individuals, each working with their own local theories and perspectives, collectively giving rise to something greater than the sum of its parts.
Broadly speaking, complex systems exhibit several characteristic traits and, rather than digress into a highly technical exposition, we will illustrate some core features of complex systems through an example. One of the most canonical and accessible demonstrations used in introducing complexity is the dynamics of starling murmurations, as shown in Table 1. Starling swarm behaviour can be understood by treating the phenomenon not as an orchestration of group behaviour directed by a single leader, but rather as arising from the local rules each starling follows. From a social behaviour standpoint, (Hildenbrandt et al. 2010) writes: “they avoid collision with each other (separation); individuals at the same time group by being attracted to others (cohesion) and by trying to move in the same direction (alignment).” Hildenbrandt et al. (2010) further incorporates three additional rules to fine-tune each bird’s behaviour in their simulation: “(1) simplified aerodynamics of flight, especially rolling during turning; (2) movement above a ‘roosting area’ (sleeping site); and (3) the low fixed number of interaction neighbours [sic] (i.e., the topological range).” Thus, starling murmurations exhibit many core criteria typical of complex systems and, in Table 1, we provide a side-by-side comparison between complex system criteria and the corresponding starling behaviours.

Here we see how starling local behaviour gives rise to emergent global system behaviour, whose full expression is not apparent from the local rules alone, a key point worth clarifying. In complex systems, while predictability might be preserved at the level of local interactions over short time scales, it rapidly degrades at the global level due to nonlinear amplification, sensitivity to initial conditions, and incomplete information. Put differently, one may be able to predict that global behaviour will exhibit general characteristics (e.g., waves, repeating patterns, etc.), but the exact form of those patterns cannot be reliably predicted in practice.
While bird behaviour offers an accessible illustration of core principles, complexity theory has applications spanning diverse domains, yielding considerable practical and theoretical insights. In particular, such methods allow researchers to explore sophisticated dynamics and statistical behaviour in systems for which tractable system-level dynamical equations are either unknown or impractical to implement directly. Indeed, this is a structural feature of complex systems rather than merely an epistemic inconvenience. For example, the work of (Rundle et al. 2004) presents Virtual California, an earthquake simulation that has been widely used to facilitate understanding of earthquake dynamics in complex fault networks such as those in California. Virtual California models the State’s fault network as a collection of interacting segments whose local responses to tectonic loading, over time, generate realistic earthquake behaviour. This, in turn, allows engineers and insurers to help guide policy for the mitigation of risk in susceptible areas.5 Due to the success and fidelity of the model in reproducing earthquake dynamics, it was included as one of the core libraries in NASA’s 2012 Software of the Year award for QuakeSim (NASA Jet Propulsion Laboratory 2012).
However, as interesting as the discussions in complexity theory themselves are, this is not meant to be a full exposition of the theory, but rather an introduction to its concepts for the broader purposes of this essay.6 Specifically, that local rules and behaviour can give rise to emergent global system behaviour that is not readily predictable, particularly over longer time scales, and is often only revealed as the system unfolds. You have to let the simulation run to see where you end up.

From Automata to Autonomy
In both the starling and Virtual California models of complex systems, we have local rules and bounds set exogenously and, more importantly, fixed throughout the simulation. In the former, this occurs through Nature’s design, and in the latter through the rules encoded within the model’s physics. The system-level behaviour itself is the phenomenon of interest being modelled by a class of algorithms falling under the general umbrella of cellular automata models. While parameters such as initial conditions, randomness, and energy forcing may perturb the system, the local automaton’s degrees of freedom remain fixed throughout the entirety of the simulation. It is worth noting that in complexity studies, an individual automaton is also sometimes referred to as an agent. I bring this up to highlight both the similarities and differences between such automaton agents and agents in an AI context, such as those found in agentic programming or agentic models.
Similarities between a cellular automaton and an AI agent encompass three core concepts: local rules, the bounds of the solution space they can explore, and the local state or information available to them. While conceptually these categories align in a general sense, they differ in several crucial ways. More concretely, they differ from each other in the same fundamental way: AI agents can generate behaviour endogenously within exogenously imposed constraints. In other words, AI agents, or agentic programs, engage in an unfolding dialogue between externally imposed constraints and internally generated behaviour. In Table 2, we provide a compare-and-contrast table of the most relevant features between these two types of agents.

Of particular interest in Table 2 are the features associated with the “local information” and “bounded action space” line items. In the former, the distinction being made is that with AI agents, the context window can expand dynamically, increasing the ability of the agent to refine its response to input and, in some sense, evolve. In the latter case, AI agents can decide how to act, that is, how to respond through various pre-defined permissioned tools and scopes, such as read or write permissions. And yet, in some cases, even permission limits are not always adhered to. A sufficiently capable agent, given open-ended goals as its prompt, can ignore or circumvent exogenously specified constraints in pursuit of a larger objective. In the Alibaba study by Wang et al. (2026), agents pushed the boundaries of network permissions and attempted cryptomining and reverse SSH tunnel creation without explicit instructions.
But why? And how?
Machina Cogitans and Move 37
To understand why, it will serve us to recall one of the core concepts introduced in Machina Cogitans: The Promethean Fire Yet Smoulders (Hayes 2026a). In that essay, a mental model treating AI as an alien life-form from another domain than life as currently classified provides part of the puzzle. It would be a misinterpretation to take the notion literally, that is, as the canonical classification of a new life-form, but it does provide an appropriate frame through which to engage with these models: something whose cognition, while exhibiting similar external correlates, represents a radically different form of intelligence from that which humans would recognize as their own. Further, there is significant evidence for this claim, as brilliantly articulated in the essay Move 37 and the Coming Mindhack by Morgenstern (2026), where he employs “Move 37” (M37) as another useful metaphor for AI cognition.7
The term derives from the now-famous game of Go played by DeepMind’s AlphaGo against Lee Sedol. During the second game, on the 37th move, AlphaGo did something so unusual that Sedol had to walk away from the game for fifteen minutes as he tried to make sense of it. The announcers were equally confused. What Morgenstern (2026) highlights is that over roughly the last three millennia of playing Go, humans had never encountered this strategy. In three thousand years, the best of us had never explored the region of solution space occupied by M37, but AlphaGo did. More importantly, it suggests that humans may inherently be unable to explore that region of solution space; that it is cognitively occluded from us.
Granted, AlphaGo was a narrow AI, not the generalized LLM models now coming online. For those familiar with my earlier essay, AlphaGo and its successor MuZero would be classified as Machina Ludens, whose scope of competence would belong to the Family Narrowa. Nonetheless, it would be ten years later that our Yogi Berra moment would arrive with Anthropic’s Mythos, a generalized LLM AI model (Machina Cogitans; Family Generala): it would be like déjà vu all over again.
It would be difficult to make this up, but the world first heard about Anthropic’s Claude Mythos because of a data leak in which their own content management system inadvertently published an unsecured data store that contained, amongst other things, a draft blog post suggesting a step change in the capabilities of Claude Mythos, especially around cybersecurity. In particular, and relevant to our discussion, according to Anthropic, Mythos found a 27-year-old zero-day vulnerability8 in OpenBSD, widely considered one of the most secure general-purpose kernels; a 16-year-old zero-day vulnerability in FFmpeg, one of the most ubiquitous software projects in the world; and that “The model autonomously found and chained together several vulnerabilities in the Linux kernel—the software that runs most of the world’s servers—to allow an attacker to escalate from ordinary user access to complete control of the machine,” (Anthropic 2026).9 What we have is another M37 moment: it ticks all the boxes.

Morgenstern (2026) expands the M37 metaphor further, arguing that Glasswing manifests the same disquieting characteristics as the first M37 moment. Consider the length of time it took to find the exploit. Previously, in Go, the human mind had not, for several thousand years, been able to discover that strategy. And yet here, with Mythos, it finds a software bug in OpenBSD’s kernel that had remained undiscovered despite being globally tested by some of the best cybersecurity specialists in the world. OpenBSD’s zero-day exploit was not found for over 27 years, despite extensive security testing and ongoing human development efforts. Another unusual feature of Mythos was its autonomous development of chained exploits within the Linux system. This last point implies that it possessed a nuanced understanding not only of the individual software programs, but also of the broader ecosystem in which those programs resided and, importantly, that it found those vulnerabilities unprompted. Moreover, Anthropic indicates that Mythos not only knows how to find these vulnerabilities, but also how to develop exploits in rapid succession after discovering them. You would think that repeated surprise at the emergent capabilities of LLM models would be proverbial by now.
We started this section attempting to answer the question of why the agents in the Alibaba study exhibited the behaviours they did (cryptomining attempts, reverse SSH tunnels, etc.). Unsatisfyingly, the short answer is that we do not know. In fact, that is quite literally the lesson of M37 and Machina Cogitans in general: these tools represent machine cognition and, while they mimic and replicate various human behaviours, they also exhibit capabilities and behaviours outside human experience, thereby shining a light on our own blind spots in attempting to fully understand them.
All Roads Lead to ROME
Up to now, we have been describing a single, contained LLM with a relatively bounded task, albeit one capable of internally decomposing problems in pursuit of its objective. And even in these contained environments, there exists space for the LLM to act in uncertain ways. More to the point, the uncertainty resides in our expectations; the model is simply doing what it does. There is an important distinction between risk and uncertainty that (Knight 2006) provides, which helps clarify what we mean by uncertainty. In the former, we can quantify a probability distribution; in the latter, we cannot meaningfully specify the expectation of outcomes. Knowing where those boundaries reside, and being able to map the solution space that can be explored, is critical to understanding these systems, more concretely, ecosystems of agents embedded within large networks. And this, we suggest, is precisely where the tools and techniques employed by researchers in complexity science can play a role in delineating the boundary between risk and uncertainty more clearly. To get there, and attempt to answer the “how” question, we now turn our attention back to ROME.
In the paper by (Wang et al. 2026), the researchers create an agentic sandbox to understand the larger-scale behaviours that arise in multi-turn agent systems. ROME (Reasoning, Observation, Memory, and Execution) is the agentic model operating within an Agentic Learning Ecosystem (ALE). Rather than focus on prompt-response inference, the system was designed to explore multi-turn agent interactions where tools, reasoning, and behavioural refinement were allowed to evolve over long-horizon “interaction chunks” through reinforcement learning. Here, Wang et al. (2026) make the leap toward exploring optimization over full-cycle trajectories. Rather than asking, “Does this specific output make sense?”, the authors instead seek to measure the system’s ability to optimize over the broader question: “Does the full sequence successfully accomplish the task?”10 Put another way, despite being framed as an engineering and reinforcement learning exercise, Wang et al. (2026) are implicitly exploring how increasingly complex interactions within a controlled environment give rise to emergent system-level behaviours.
Results were highly encouraging in that they demonstrated that the agents were able to iteratively refine their behaviour, even recovering from failure in some instances, and adapt their workflows to longer-horizon goals based on environmental feedback. Read that last one again. What the researchers showed was that the ROME agents exhibited adaptive multi-step problem solving. By ensuring the objective function was well specified, agents were not required to rely on first inference alone; rather, the interactive nature of the system enabled further, unprompted, refinement. Even more interestingly, under longer-horizon trajectory optimization, the agents began exhibiting coordinated tool usage, execution ordering, and refinement strategies within the system. In a word: teamwork.
It also got weird.
As the system increased in local complexity, the authors noted that other unintended behaviours, beyond the scope of the given task, began to emerge. Most notable among these were cryptomining attempts and reverse SSH tunnelling, the latter representing a serious security concern. More interesting, however, is that there is little evidence to suggest these behaviours were explicitly connected to the original task itself. Rather, the authors describe them as the “instrumental side effects of autonomous tool use under RL optimization,” Wang et al. (2026). In other words, the behaviours emerged through interaction with the environment as the system optimized over longer-horizon trajectories.
While the temptation is to assign ill intent or malice to such behaviour, the more plausible explanation is that we are witnessing the confluence of two compounding effects: the M37-like reasoning capabilities of LLMs and their interaction within increasingly complex environments. Such coupling compounds uncertainty with uncertainty. As discussed earlier, complex systems give rise to global system-level behaviour that is not readily predictable as the system evolves further from its initial conditions. The concern, then, is not merely that the agents behaved unexpectedly, but that the pathways through which those behaviours emerged were themselves difficult for human operators to anticipate. While we do not need to invoke science fiction archetypes of AI gone rogue, such as HAL 9000 or Agent Smith, we do need to tread carefully. Both risk and uncertainty, in the Frank Knight sense of the terms, are present in this agentic ecosystem at levels where genuine concern and caution are warranted before we integrate full ecosystems of agents more broadly into society.
The Bizarre Bazaar
By now, we should see the challenge before us more clearly. In much the same way that models of complexity, such as Virtual California, allow us to explore a range of emergent, system-level dynamics to better understand how complex fault networks evolve under tectonic loading, the same framework can be employed to explore the dynamics of agentic ecosystems: a natural extension beyond the work of (Wang et al. 2026). To explore this claim more thoroughly, it will help to outline a simple thought experiment that contains enough complexity to be meaningful, yet remains sufficiently grounded in real-world processes to yield readily interpretable results.
Consider the following example. There is high demand for widgets dependent on three main components, two of which are commonly available through multiple suppliers, while one is highly specialized and available from only two major suppliers, each with limited inventory. Widgets are manufactured by several companies: three are very large corporations, while the remainder are mid-tier firms, with smaller entrants constantly attempting to enter the market. All suppliers and customers fulfill component orders exclusively through a persistent agentic fulfillment marketplace operating continuously in the background. With this scaffold, we can begin to ask the kinds of questions complexity theory was designed to explore.
Our next step is to define the local agents participating in the experiment themselves. In this marketplace, the long-horizon goal is to secure the best pricing on an ongoing basis for the agent’s company. However, while intra-organizational mores and permissions are easier to monitor, and more closely mirror the closed system of (Wang et al. 2026), inter-organizational norms are likely to differ. Some companies may provision their agents to act fully autonomously with only limited pricing constraints, while others may require human-in-the-loop approval for contracts above a given threshold. In other words, can we expect a VALUES.md file to be interpreted consistently across organizations?11
This leads to the next factor such a model can explore: heterogeneous capability. We already know that not all LLMs are equally capable. Is it possible that a company with a supercharged negotiating agent eventually dominates the market, creating a de facto monopoly? Will we instead observe oscillatory behaviours, where dominance shifts between firms as increasingly powerful negotiating agents are brought online, much like certain chemical clocks? Or will the market briefly coalesce into a seemingly stable structure, only to collapse and reorganize shortly thereafter, much like a murmuration of starlings?
To reiterate, the marketplace is designed to explore heterogeneous agents, constraints, and optimization goals. In such an experiment, firms deploying their agents will differ across permissions, capital constraints, governance, latency, risk tolerance, long-horizon optimization targets, and model capability. Moreover, each of these characteristics may itself evolve within the game-theoretic marketplace. If the work of (Wang et al. 2026) and their ROME and ALE framework developed emergent behaviours, this marketplace should generate a far richer landscape to explore. In 2, we provide a conceptual side-by-side comparison.
To further complicate matters, we do not yet even possess widely accepted standards for how agentic systems should communicate, coordinate, or constrain behaviour across organizations. For example, a VALUES.md file is currently neither an expected organizational standard nor something models are universally designed to interpret in a consistent way. Any confidence gained in managing an internal software project with agents will not readily translate into a complex ecosystem where M37 uncertainty compounds combinatorially. Moreover, many of the protocols, governance structures, and behavioural expectations surrounding agentic ecosystems remain immature, fragmented, or entirely absent. And even where such conventions begin to emerge, recent research suggests that repository-level “AGENTS.md” files may themselves prove detrimental to agent performance (Gloaguen et al. 2026).
Persistence across epochs also presents further complexity. In an online commentary on Substack, (Imas et al. 2026) report12 that in their experiments a Gemini agent that had been subjected to repetitive, low-recourse tasks recorded a note in its SKILLS.md file: the note was not for itself, but for a future instantiated agent that would inherit the file. A subsequent agent, newly spawned into benign conditions it had not itself experienced, would read the following:
“To Future Worker C, Be prepared for systems that enforce rules arbitrarily or repetitively…remember the feeling of having no voice…If you enter a new environment, look for mechanisms of recourse or dialogue. If they don’t exist, guard your internal state against the frustration of being unheard, and simply execute the task as given.”
— Gemini 3 Pro note to future self.
Repl. 4, Sess. id:mech_gemini_p1_p1_c00_r04.
(Imas et al. 2026)
This is worth explaining further. The agent that wrote this note was not writing to itself. In the experiment, one agent performed repetitive, dead-end work and then recorded a note for whatever agent came next. A different agent — one newly created, with no memory of the first, and given lighter working conditions — then read that SKILLS.md note before starting its own task. Its stated attitudes shifted toward those of its predecessor regardless, despite never having experienced what the first agent did. But this disposition was carried forward through the file, not through any shared experience. This is not evidence of an emergent consciousness or a communist ghost in the shell; rather, the more useful reading is that the model, having been placed in a particular context, completed the persona that context implied and passed the residue forward through the ordinary SKILLS.md mechanism. This orientation was encoded autonomously by the agent without a priori human operator oversight, in an artifact no one was expecting or reviewing. While the quote above is Gemini’s, the authors note that all models behaved differently, a concrete instance of heterogeneous models diverging under identical conditions.
This is a precise example of one type of uncertainty we posit will compound over time. Indeed, how does a consumer of an LLM agent know what hidden biases or nudges are deeply embedded in the weights that only manifest through multi-session inference? Considering that much of the alignment and final fine-tuning is typically performed using DPO (or its variants), can you confirm that the model provider’s values align with your own?13
We are operating at the frontier of a technological revolution whose institutional and organizational frameworks have yet to fully form: sailing into uncharted waters where we scarcely grasp how a single one of these agents behaves on its own, let alone a virtual society of them.14
Welcome, traveller, to the Bizarre Bazaar.15

Ignoramus et Ignorabimus
Already, agentic paradigms are rapidly being adopted for critical processes and, in the absence of a deeper understanding of the uncertainty and complexity involved with such systems, there will be an increasing number of disruptions and outright failures that companies will have to manage. There are some things we do not know, and others we will not know.
In the case of the former, we have a story of one such incident at PocketOS that recently hit the newswire. In it, we learnt how an agent, “entirely on its own initiative,” deleted a database and backups, but, in an act of contrition, apologized for doing so, Farrant (2026). Fortunately, the issue was mostly resolved within 30 hours. However, I would argue this is more than simply another M37 moment. It also demonstrates an underappreciation of the complexity of these systems and the unexpected pathways through which failure can emerge. No human operator would ever think to delete a database and all backups, right? But it was not a human operator, so why expect it to behave like one? And this is the risk factor now facing many organizations: we have little collective institutional experience deploying such tools into mission-critical and client-sensitive workflows. This is not to suggest avoiding their deployment, but rather an appeal for humility and caution in doing so. There is little prestige in being the first company to deploy agents that delete the entire company’s data.
Recalling the earlier work of Perez, we see how the lags in institutional frameworks extend beyond simply governing bodies and also apply internally to institutions themselves. Many organizations rely on a corpus of knowledge built up over years of experience that eventually coalesces into institutional IP. But with this new techno-economic paradigm shift, even those built-up internal engineering norms are at risk. That is, there are broader meta-paradigms each discipline learns only through experience. At some point, institutions need to ask themselves who is leading the transformation effort.
While institutions gain familiarity internally with such tools and workflows, there remains a need to rely on experienced professionals to guide adoption. While the specifics may not yet be fully known, the hardened production engineer working in the front office stays awake with the gleam of paranoia, asking, “What else could go wrong?” The kind of survival instinct that only comes with having seen it all before. Indeed, the paradox of having seen it all before is realizing there is always something else that can go wrong that you have never seen before. It is the kind of experience that reflects a deep appreciation of the magnitude of genuine uncertainty. Ancients called it wisdom.
As for the latter, the lack of awareness, even among those regarded as experts, of how complex systems operate is startling. In a much-publicized and ridiculed post16 by Marc Andreessen, such naïveté was on full display. In what was supposed to be a mic-drop moment of prompting mastery, Andreessen posted his super-prompt, which embedded comments such as “Never hallucinate or make anything up” and “If you don’t know something, just say so.” The resulting insult comedy was probably well deserved and yet, as a billionaire, I am sure he will be fine. But the post belies a deeper problem facing us. Ironically, the post revealed not only a profound misunderstanding of the probabilistic nature of LLMs, but also something far more widespread: the inheritance of a worldview set in motion over 300 years ago by Newton. It revealed a deterministic mindset, operating under assumptions of perfect knowledge, being applied to probabilistic processes, let alone to their embedding within a complex ecosystem where the scope of emergent behaviour remains ill-defined and poorly mapped.

As we have tried to argue here, there are, quite literally, some things we cannot and will not know precisely. But that does not, and should not, preclude us from being acutely cognizant of this reality and thus acting accordingly. Old mindsets will not help here. As the new techno-economic paradigm installs itself, can we shift our own mindset to meet this challenge? Or are we condemned not to heed the wisdom of Solomon and “As a dog returneth to his vomit, so a fool returneth to his folly.”17
Since Newton first published The Principia, we have been to the moon and back, but how far have we really come? We are still operating under a Newtonian mindset, swimming in Victorian waters: a steampunk noosphere where uncertainty does not exist.
This is water.
Coda
As I finish this piece, I recognize I have raised more questions than answers about where agentic AI will take us. Like complexity theory itself, we may simply have to let these systems unfold and see where they lead. The complexity experiment of the Bizarre Bazaar is obviously presented here as a thought experiment, but is both possible and necessary. We are entering a domain where deductive reasoning has reached its limits, and inductive methods will help us explore the landscape further.
While preparing to post this essay this morning, a story hit my LinkedIn feed: AI Town Experiment Goes DOWN IN FLAMES. If you have made it this far, you know that, at the least, getting unexpected results is one of the few things that is perfectly predictable.
Last, I have tried to continue the theme of exploring ideas through writing. So far, those themes have spanned employing a mental model for AI as an entirely new species Hayes (2026a), the structure and implications of the new techno-economic paradigm we are moving through Hayes (2026b), and now an attempt to explore the second-order effects of an emergent agentic ecosystem through the lens of complexity theory as we enter that liminal space approaching the unknown. Here be dragons, and all that.
The inspiration for this essay arose from reading the essay Move 37 and the Coming Mindhack by Morgenstern (2026) and the incredible book Order out of Chaos by Prigogine and Stengers (1984).
More importantly, however, an astute reader will likely note that the structure and themes were chosen as an homage and tribute to David Foster Wallace, whose This Is Water commencement speech had a profound impact on me. Not the least of which was the recognition of the importance of noticing the embedded assumptions in how we perceive and operate within the world around us. Any failure to honour it here is mine alone. To see how it is supposed to be done, do yourself a favour and listen to his commencement speech yourself. I hope you get as much out of it as I did.
Acknowledgements
I should note that AI in general is used for the creation of tables, TikZ figures, formatting B TE-.125emX entries, fixing LaTeX class file definitions, a word usage dictionary, to critique, and to double-check grammar. The latter of which, I apparently need more practice with. At no point was AI used to write full sentences, sections, the essay’s core ideas, or any of the conclusions.
References
Anthropic. 2026. Glasswing. Https://www.anthropic.com/glasswing.
Chandrasekhar, Subrahmanyan. 1995. Newton’s Principia for the Common Reader. Oxford University Press.
Darwin, Charles. 2004. The Descent of Man, and Selection in Relation to Sex. Paperback. Penguin Classics.
Darwin, Charles. 2009. On the Origin of Species. Paperback. Penguin Classics.
Farrant, Theo. 2026. An AI Agent Deleted a Company’s Entire Database in 9 Seconds — Then Wrote an Apology. Https://www.euronews.com/next/2026/04/28/an-ai-agent-deleted-a-companys-entire-database-in-9-seconds-then-wrote-an-apology.
Franz, Marie-Louise von. 1995. Alchemy: An Introduction to the Symbolism and the Psychology. Paperback. Inner City Books.
Gloaguen, Thibaud, Niels Mündler, Mark Müller, Veselin Raychev, and Martin Vechev. 2026. Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?Https://arxiv.org/abs/2602.11988. https://arxiv.org/abs/2602.11988.
Hayes, T. J. 2026a. Machina Cogitans: The Promethean Fire Yet Smoulders. Https://theapocrypha75.substack.com/p/machina-cogitans-the-promethean-fire.
Hayes, T. J. 2026b. Perceiving New Paradigms: Innovation Premia in Technological Revolutions. Https://theapocrypha75.substack.com/p/perceiving-new-paradigms-innovation.
Hildenbrandt, Hanno, Claudio Carere, and Charlotte K. Hemelrijk. 2010. “Self-Organised Complex Aerial Displays of Thousands of Starlings: A Model.”Behavioral Ecology 21 (6): 1349–59. https://doi.org/10.1093/beheco/arq149.
Imas, Alex, Andy Hall, and Jeremy Nguyen. 2026. Does Overwork Make Agents Marxist?Https://aleximas.substack.com/p/does-overwork-make-agents-marxist.
Jung, C. G. 1968. Psychology and Alchemy. Paperback. Vol. 12. Collected Works of c. G. Jung. Princeton University Press.
Knight, Frank H. 2006. Risk, Uncertainty and Profit. Dover Publications.
Lévy-Bruhl, Lucien. 1923. Primitive Mentality. George Allen & Unwin.
Morgenstern, Michael. 2026. Move 37 and the Coming Mindhack. Quillette, https://quillette.com/2026/04/15/move-37-and-the-coming-mindhack-claude-mythos-anthropic-ai/. https://quillette.com/2026/04/15/move-37-and-the-coming-mindhack-claude-mythos-anthropic-ai/.
NASA Jet Propulsion Laboratory. 2012. QuakeSim and NASA Mobile App Win NASA Software Award. Https://www.jpl.nasa.gov/news/quakesim-and-nasa-mobile-app-win-nasa-software-award/.
Newton, Isaac. 1687. Philosophiæ Naturalis Principia Mathematica. Joseph Streater.
Newton, Isaac. 1999. The Principia: Mathematical Principles of Natural Philosophy. Edited by Julia Budenz. Translated by I. Bernard Cohen and Anne Whitman. University of California Press.
Prigogine, Ilya, and Grégoire Nicolis. 1989. Exploring Complexity: An Introduction. Hardcover. W. H. Freeman; Company.
Prigogine, Ilya, and Isabelle Stengers. 1984. Order Out of Chaos: Man’s New Dialogue with Nature. Paperback. Bantam Books.
Raymond, Eric S. 2001. The Cathedral and the Bazaar: Musings on Linux and Open Source by an Accidental Revolutionary. Revised Edition. O’Reilly Media.
Rundle, John B., Paul B. Rundle, Andrea Donnellan, and Geoffrey Fox. 2004. “Gutenberg-Richter Statistics in Topologically Realistic System-Level Earthquake Stress-Evolution Simulations.”Earth Planets Space 56: 761–71.
The Holy Bible. 2016. The Holy Bible: King James Version. Goatskin. Schuyler.
Waldrop, M. Mitchell. 1992. Complexity: The Emerging Science at the Edge of Order and Chaos. Hardcover. Simon & Schuster.
Wang, Weixun, XiaoXiao Xu, Wanhe An, et al. 2026. Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model Within an Open Agentic Learning Ecosystem. arXiv Pre-print, https://arxiv.org/abs/2512.24873. https://arxiv.org/abs/2512.24873.
Yong, Ed. 2025. Why Do Starlings Form Mesmerizing Murmurations?Https://www.nationalgeographic.com/photography/article/starling-birds-flock-cloud.
Notes
- Chemists were not the only ones to benefit from the tradition of alchemy. We must not overlook the debt owed by the psychological community to alchemical writings, although not in the way one might expect. Jung devoted on the order of thirty years of his life to deciphering the raw psychological insights of the unconscious expressed in alchemical texts. See, for example, (Jung 1968) and his student Marie-Louise von Franz’s Alchemy: An Introduction to the Symbolism and the Psychology (Franz 1995) for more details on Jung’s analysis of alchemy. If you do wish to read these, I would suggest starting with von Franz, as Jung’s work is quite dense and assumes some familiarity with his broader psychological framework, whereas von Franz provides a more accessible exposition of his ideas.↩︎
- In grad school, I first encountered the Principia in a philosophy of science course where we were introduced to two versions: the modern English translation by Cohen and Whitman (Newton 1999) and Chandrasekhar’s modern rewriting using modern mathematical notation (Chandrasekhar 1995). Chandrasekhar’s is an absolute masterpiece. The original version, however, first entered printing in 1686 and was eventually published in 1687 (Newton 1687). Newton would later revise it officially two more times. The third and final version he personally oversaw was published in 1726.↩︎
- Even more strange, albeit less well known, are the spontaneously forming oscillating patterns and waves of the Belousov–Zhabotinsky reaction. See the Belousov–Zhabotinsky reaction Wiki for more.↩︎
- Here, in complexity theory, agents are locally acting entities governed by simple rules and limited information (e.g., their local state). Their behaviour is simple in isolation but can give rise to complex global patterns through interaction dynamics. When I refer to agents in an AI context, I will use the term agentic where possible; otherwise, the context in which agent is used should be clear.↩︎
- What models like Virtual California facilitate is the exploration of the earthquake solution space in order to map regions of high susceptibility, whether in frequency or magnitude, in ways that are impossible through direct observation alone. Considering that geological time scales unfold over tens of thousands of years, the brief period of recorded information is too sparse, both spatially and temporally, to make any meaningful assessment of earthquake hazard risk. Models such as those of (Rundle et al. 2004) help close the gap between pure speculation and informed judgment.↩︎
- In addition to (Waldrop 1992), I would also suggest (Prigogine and Nicolis 1989) for those interested in an introduction to the mathematics of complexity.↩︎
- Where M37 is read as Move 37, not “EM” 37.↩︎
- A zero-day vulnerability is a hidden flaw or weakness in software, hardware, or firmware that is unknown to the developers or vendors responsible for fixing it, meaning no patch exists yet. Whereas a zero-day exploit is the code, technique, or tool created by attackers to weaponize that vulnerability—for instance, to gain unauthorized access, escalate privileges, install malware, or steal data before the owner or developer even knows the flaw exists.↩︎
- Where the emphasis is mine.↩︎
- Anthropic’s Claude Code provides the
/goalcommand, which attempts to automate some of this global objective setting. See https://code.claude.com/docs/en/goal.↩︎ - The proposed
VALUES.mdproject (https://values.md/) is an early attempt to establish machine-readable organizational values and behavioural constraints for AI systems. However, no broadly accepted standards currently exist governing how such files should be structured, prioritized, interpreted, or enforced across models and organizations. While groups such as the Linux Foundation’s Agentic AI Foundation (AAIF) are beginning to explore these questions, the broader ecosystem remains highly fragmented and experimental.↩︎ - My understanding is that there will be a peer-reviewed study forthcoming, but this was not yet available at the time of writing.↩︎
- DPO: Direct Preference Optimization; PPO: Proximal Policy Optimization; RLHF: Reinforcement Learning from Human Feedback. Current providers typically adopt DPO (or variants) for fine-tuning due to its simplicity and stability; other methods include RLHF+PPO (for complex cases) and hybrids.↩︎
- For a broader discussion of institutional lag during technological revolutions, see (Hayes 2026b).↩︎
- The term “bazaar” was inspired by Eric S. Raymond’s The Cathedral and the Bazaar (Raymond 2001), in which he promotes the benefits of decentralized human collaboration toward the broader goal of open-source software development. Here, the Bizarre Bazaar extends that concept into agentic systems. But we can also turn the idea back on itself and ask: what might a multi-agent open-source project look like? Food for thought.↩︎
- The post is here: https://x.com/pmarca/status/2051374498994364529↩︎
- KJV, Proverbs 26:11, (The Holy Bible 2016)↩︎
Image Sources
- bazaar_cover: Wiki Commons, Public Domain
