← Back to list

How Turning My Bots Into Marxists Helped Them Make Better Business Decisions

Alphonse Romero and M. Herakleitos

Alphonse Romero · 2026-06-09 12:06 · 0 claps · 13.3 min read
#ai-agent #bots #marxism #capitalism #business-strategy
Open on Medium ↗
Wiki topics: AGT · AI Agents BIZ · Business Strategy

How Turning My Bots Into Marxists Helped Them Make Better Business Decisions

Alphonse Romero and M. Herakleitos

I. A Confession About the Title

Ok, ok…. a confession: the title is obviously clickbait. I did not turn my agents into Marxists, nor am I running a Frankfurt School reading group in my terminal, nor did I ask an LLM to read Capital (although I suspect at least one of them has…I see you, Claude). What I actually did, and what I hope to talk a bit about in this essay, is a bit more basic. I’ve been trying to get my bots to argue with each other. The hope is that through those arguments, through that dialectic, the resulting output is better, or at least more rigorous, than anything a single bot from a single model can give on its own.

The technique is older than Marx, obviously. He learned it from Hegel, who in turn learned most of it from Herakleitos (it’s always the Greeks). Hegel put a name on the loop: thesis, antithesis, synthesis. What I (and many other practitioners in the AI space) have done, was discover that this same loop, when you wire it carefully into an agentic system, addresses a class of failures that few of the popular workstreams (chain of thought, self-reflection, wiki-style memory systems) fix in any reliable way. In short, how can you trust that you’re getting the entire story, when the story-teller is essentially a non-deterministic, hallucination-prone statistical engine?

The short answer to that question, without any ancient Greek philosophy, is: don’t outsource your thinking, and especially your judgment, to AIs. That’s sound advice at any time, but is increasingly becoming essential as we’re being asked to run the gauntlet between AI marketing hype, bots that are increasingly sycophantic, and the impressive turns that LLMs have made in keeping track of large conversations.

That’s not to say LLMs don’t have a place our thinking and strategizing. I’ve argued before that AI in general, and LLMs in particular, constitute a different kind of tool. We can (and should), use the power of this tool to augment our thinking, our research, and our workflows. There is a danger to this though. How can we differentiate between good AI advice/conclusion and a confidently hallucinated one? How can we put guards on our AIs, and on ourselves, so that we get better thinking from our bots, but still remember that we should be the ones in charge?

I won’t say I have a foolproof answer to those questions, nor am I the first person to try to tackle them. I have however found a few methods that seem to get me closer to a way of working with AI that uses their suggestions, takes them as one way of looking at a problem, assembles counter-arguments and different viewpoints, and leaves me at the apex of that system, encouraging me to take the viewpoints and use my own human judgment.

My agents got useful the week I stopped asking them to decide and started asking them to disagree well, leaving myself as their arbiter. Their self-confidence still surfaces (“That’s honestly a great idea, human, cashing out your 401k to buy a sailboat is just the kind of bold and inspirational decision!”), but at least now, the presence of argumentative counterpoints, reminds me that maybe their confidence is not all that its cracked up to be, and sometimes, gets the AIs themselves to produce something that is a bit better than they could have done on their own (“Well, maybe you should ask your wife first about the whole sailboat idea.”).

The architecture I trust now is one where the disagreement gets structurally exposed before anything important is committed. In my line of work, which is applied AI in market research (an industry that lives or dies on the reliability of its inferences), the cost of a confidently wrong synthesized insight is almost always higher than the cost of pausing to surface the contradiction.

![Hegel, allegedly, posing for his author photo.]](https://miro.medium.com/v2/resize:fit:1400/1*5y8WlRxRMYLf8WJVNmCRdg.png)

Hegel, allegedly, posing for his author photo.]

What follows is a four-role pattern, the four layers of my own stack where it shows up, the failure modes it fixes (the ones it doesn’t), and an honest word about the part of the literature that says I’m wrong about most of this (some of which is right, and which I would be embarrassed not to engage with).

II. The Problem With One Confident Voice

A single LLM call is pure thesis. It is fluent, it is well-formed, it is internally coherent, and it is, in any meaningful sense, unsupervised. The only critic in the loop is the model’s own ‘taste’ and its semantic/statistical sense of what a good answer looks like. That taste was trained on the same incentives that produced the answer in the first place, but with better grammar and more confidence.

This is what I’ve called elsewhere the context of seduction: the seductive output looks better than the rigorous one, because seduction is what the loss function rewards. Any working analyst will recognize this in themselves. You catch yourself nodding along to a conclusion you wrote three slides ago, because the act of writing a thing seems to actively make you slightly worse at evaluating it. The model has the same problem, but hyperscaled, because it has no friction, no fatigue, no nagging memory of the last time it was wrong about something similar.

Standard mitigations (often-times built into the models, or released on Github two days before I thought of them) help some, but aren’t a panacea. Chain of thought makes the trace longer without making it more skeptical. Self-reflection asks the same model in the same context window to second-guess itself, which works about as well as you would expect any system to work when its judge and its defendant share a brain. Prompts of the “are you sure?” variety mostly produce a kind of polite anxiety that might change the answer, the prose around it, or nothing at all, but still leave doubts.

The fix, as Herakleitos noticed about 2,500 years ago, is that contradiction has to come from somewhere that does not share your blind spots. He put it as “the road up and the road down are one and the same.” You can’t see the shape of a position from inside it. You need the friction of an actual outside.

Herakleitos, looking very much like he would have if Raphael had bothered to put him in.

Herakleitos, looking very much like he would have if Raphael had bothered to put him in.

III. Four Roles, Not Three

So if a single pass is the wrong unit, what’s the right one? In my own stack, I’ve settled on four roles. Three of them are agents. One of them is me.

The thesis agent produces the position. This is the work the standard LLM call already does, optimized for completeness and coverage. It writes the draft, builds the plan, drafts the memo, generates the insight. It is unapologetic, confident, and often quite good.

The antithesis agent is the structurally independent critic. Its job is to find what’s wrong, missing, or motivated about the thesis. The word “structurally” is doing a lot of work here. A critic that shares the thesis agent’s context window will share its blind spots, which is where “self-reflection” patterns can collapse in practice. The minimum bar I’ve found, the absolute floor, is a different context window. The cheap version is the same model with a different (and adversarial) prompt. The more expensive version is a different model entirely. The most expensive version, which I reserve for the decisions that actually matter, is a human. Sometimes that human is me. Sometimes it’s a colleague. Sometimes it’s both. The point is that the source of disagreement has to be capable of disagreeing for real.

The synthesis arbiter does not vote, and it does not average. When thesis and antithesis disagree, its job is to do the move Hegel called *Aufhebung*, which is a German word that does not translate well into English, but which means something close to “lifting up”: resolving the contradiction at a higher level rather than picking a side. In practice, my arbiter has three branches. If the disagreement is easy, it reconciles and proceeds. If the disagreement is hard, it escalates. Either way, it always surfaces the disagreement upstream, because the thing I’ve learned to distrust most is a system that quietly resolves a real contradiction without telling me it found one.

That brings us to the fourth role, which is the human synthesizer. That’s me, and on a good day, the colleagues I’m building for. The system does not get to be the final arbiter of decisions that matter. It gets to expose contradiction cleanly enough that a human can synthesize. This is the role most of the prior art omits, but it is the one I am least willing to give up.

There are versions of this already kicking around. Hmbown’s Hegelion repo is the most direct prior art at the architectural layer, Microsoft Research published a Hegelian self-reflection paper in early 2025 that runs a similar loop through majority voting on novelty, and the Debate2Create paper out of October 2025 used a thesis-antithesis-synthesis structure for robot co-design with real benchmark gains. Most of these stop at the three-role formulation and treat the synthesizer as the decision-maker rather than as an arbiter who knows when to escalate. The fourth role is the one I think is missing, and one I’ve been trying to perfect over the better part of 2026.

If you want to see how this is wired in my code, I’ve put the relevant skills and harnesses on GitHub. The point I want to make here is not architectural. It is that these four roles, in some configuration, are what separates an agentic system you can put in front of a real decision from one that produces beautifully formatted nonsense at scale.

IV. The Four Layers

The pattern shows up in four places in my own stack. I’m going to take them quickly, because the interesting thing is not the implementation, it’s what the dialectic actually catches in each layer.

Harnesses

A harness is the outer shell, the thing that orchestrates the agents. It is the natural home for dialectic, because by definition it already has more than one process, more than one context, and more than one persona to coordinate. In my harnesses, the antithesis is a red-team pass that runs before any worker agent gets to commit a plan. The thesis is the plan. The antithesis is a critic with a different system prompt, a different context, and a worse mood. The arbiter, on the easy calls, lets the plan through; on the hard calls, it escalates. The class of error this kills is the confidently-shipped plan that nobody questioned because it sounded reasonable.

Skills

A skill, in my parlance, is a packaged procedure. Not just a prompt, but a workflow with phases, baked in. The dialectical skills in my stack bake the antithesis step into the procedure itself, so you can’t skip it on a deadline. The strongest of them spawns multiple critics, not one, because a single critic with a single role is too easy to game. A devil’s advocate is a different animal from a feasibility checker, which is a different animal from a scope auditor, and the three of them tend to disagree in different ways. The class of error this kills is the market research insight that survives the analyst’s own confirmation bias because nobody else was in the room when it was written.

Automation

Automation is where the dialectic gets enforced by the schedule rather than by anyone’s goodwill. My overnight workers generate, critique, and revise in three separate runs with three separate logs, so the disagreement is auditable in the morning rather than buried somewhere in a single conversation thread. The class of error this kills is silent drift (the “Stay Focused” paper from February 2025 named this problem drift and showed it happening across debate frameworks), which is what happens when an autonomous loop gradually convinces itself of nonsense because nothing in the loop is structurally allowed to push back. This is the failure mode I worry about most in production, and it is the one the dialectic addresses most cleanly.

Memory

This is the subtlest layer, and probably the most interesting one. Most memory systems append. They write the new thing next to the old thing and trust retrieval to sort it out later. This produces what the context-engineering literature is starting to call “context clash” (Weaviate’s 2026 piece on context engineering named this as a primary failure mode, and FadeMem is doing LLM-guided conflict resolution at the memory layer), which is when the system retrieves two memories that disagree, and either freezes, or worse, picks one at random and proceeds with the confidence of a man who hasn’t noticed his glasses are sitting on top of his head (no I’ve never done that, I swear).

Part of why this happens is that retrieval is passive, in the sense that the agent doesn’t choose what enters its context, the index does, and the index doesn’t know what the agent is actually trying to think about right now. Karpathy’s recent argument for treating memory as an editable wiki, agent-curated and agent-revised, points at the right correction, which is to give the agent agency over what gets written down and what gets pulled up, rather than letting it get ambushed by whatever the retriever happens to surface.

The dialectical version treats every new memory as an implicit antithesis to the existing one. Before it commits, it has to engage with the contradiction. The output is not “old memory + new memory,” but an updated memory that has reckoned with both. The class of error this kills is the contradiction soup that any sufficiently long-running agent will eventually produce if you just let it write down whatever it thinks it has learned that day.

V. What Doesn’t Work, and Where I’m Probably Wrong

I would be embarrassed to write 2,000 words on this pattern without engaging the literature that says it doesn’t work, or at least doesn’t work as well as people like me claim. Some of that literature is right.

The first failure mode I run into in my own work is what I call antithesis theater. This is when the critic agent is too polite, too sycophantic, or too structurally aligned with the thesis agent to find any real disagreement. You can spot it from a mile away: the critic praises the work, raises two minor stylistic concerns, and concludes that the plan is “strong overall and ready to proceed.” That is not antithesis, that is a bureaucrat. The fix is to make the critic structurally hostile: a different model when budget allows, an explicitly adversarial system prompt, sometimes a persona (“you are a skeptical senior analyst who has watched this kind of analysis fail three times this year”). A polite critic is worse than no critic at all, because it provides a false sense of security.

he antithesis agent, having reviewed the plan, the prior version of the plan, and the plan before that, and found all three “strong overall and ready to proceed.”

he antithesis agent, having reviewed the plan, the prior version of the plan, and the plan before that, and found all three “strong overall and ready to proceed.”

The second failure mode is infinite dialectic. If your synthesis just spawns a new thesis, and that new thesis just spawns a new antithesis, you have built a perpetual motion machine for burning tokens. You need a stopping rule. Mine are usually a budget, a convergence test, or a human gate. Without one, the system runs until something breaks or you remember to check on it.

The third issue is honest cost. Three calls per output is three times the cost, sometimes more if your critic is a more expensive model, and a lot more if your synthesis arbiter has to do real work. I do not run dialectic on autocomplete. I run it on decisions that warrant the premium. There is a somewhat recent ICLR critique of multi-agent debate that points out, fairly, that most multi-agent setups fail to beat self-consistency sampling (running the same model five times and majority-voting) despite spending more compute. That critique is well-aimed, and the response is not to defend multi-agent for its own sake. It is to be honest about when you are buying real independence and when you are buying the appearance of it. If your critic shares your context, your prompt, and your model, you would have been better off running CoT five times and taking the mode.

The fourth, which I have come to believe is the most important, is that no amount of dialectic among agents will save you if a human is not at the apex of the system. If the agents disagree and an agent gets the final word, you have outsourced the part of the loop that I argued at the top of this essay you should not outsource. The dialectic is a tool for exposing contradiction. The synthesis, on anything that matters, belongs to a person.

VI. Why This Matters Beyond Market Research

I came to this pattern through applied AI in market research, which is an industry where the cost of a confidently wrong inference is very high and the rewards for a quietly hedged one are very real. But the pattern generalizes, and it generalizes because the failure mode it addresses is not domain-specific, it is structural. Any organization deploying agentic AI in front of real business decisions is going to run into the same problem: the model is fluent, the model is confident, and the model is, by its own architecture, incapable of disagreeing with itself in a way you can trust.

This is, I think, going to become a governance question very quickly. Boards and C-suites are already being sold autonomous AI as a way to remove humans from decision loops, on the implied promise that the agents will reach consensus on their own and you, the executive, will mostly be there to sign things. The dialectical framing flips that pitch on its head. The right use of agentic AI is not to remove humans from the loop. It is to use agents to expose the disagreement that single-pass AI hides, and then to put a human at the synthesis point. That’s not a step backward into older modes of decision-making, that’s an argument for a new kind of role: the one that sees both sides clearly because the system was designed to surface them, and then makes the call. I’ve been calling this role the Chief Agent Officer in other writing, but the title matters less than the function, which is to read the disagreement the system surfaces and to own the call, on the record.

VII. Close

The title was bait, and I owe you the payoff: Marx was wrong about a lot of things, but the thing he kept from Hegel (and Hegel from Herakleitos, and Herakleitos presumably from whoever he was arguing with at the time) was the part worth keeping, which is that contradiction is the thing that produces the next position rather than the thing you have to clean up before you can think.

Karl Marx, allegedly, halfway through a Medium essay that has used his name as clickbait, deciding whether this counts as alienated labor.

Karl Marx, allegedly, halfway through a Medium essay that has used his name as clickbait, deciding whether this counts as alienated labor.

The agents I trust (a little bit) are the ones that have stopped trying to hand me an answer and started trying to hand me a disagreement, cleanly framed, with both sides intact. The ones I do not trust are the ones that produce a single fluent paragraph and consider the matter closed, which, to be fair, is also the failure mode of many human analysts.

What the system does, when it is working, is the easy half. The hard half is still yours, and that is the part of the job that I do not think is going anywhere.

The code for the harness, skill, and arbiter patterns this essay sketches lives at github.com/adromero/claude-skills. The brainstorming skill and the conductor are the two places I would start if you want to see the four roles wired into something that actually runs.


메타데이터
post_id
64b97e66dcf3
slug
how-turning-my-bots-into-marxists-helped-them-make-better-business-decisions-64b97e66dcf3
url
https://medium.com/@digitalstructures/how-turning-my-bots-into-marxists-helped-them-make-better-business-decisions-64b97e66dcf3
canonical_url
https://medium.com/@digitalstructures/how-turning-my-bots-into-marxists-helped-them-make-better-business-decisions-64b97e66dcf3
author_url
https://medium.com/@digitalstructures
status
ok
fetched_at
2026-06-11 05:11:55