Skip to content

A thesis on Work and Research with agents, in 2026 and beyond.

Why the future of knowledge work is an alloyed composite of un-correlated intelligences, and how to rescue agents from producing slop whilst tackling research bottlenecks.

Back to writing
TJ Chalokia
AI Engineer and Consultant

Questions

Let's ask ourselves a few questions.

Why is it, that we're not at all surprised by agents being somewhat disappointing, at fully automating knowledge work, yet? Genuinely, why is it both intuitive and unsurprising, that despite benchmarks, despite the apocalyptic marketing, the simple fact remains that agents, even long-running agents, even really reallly scawy agents named things like ""Mythos""" do rather poorly on any ambitious, and unverifiable work END-TO-END?

And why--this not a rhetorical question btw--do we STILL need humans steering and reviewing an agent's work, not everything, not every token, but some batch of work at a higher synthesised abstraction (markdown files), despite the impressive METR curves, which definitely won't be mis-interpreted by you and your boss.

The answer to those questions, at some intuitive level is obvious. Many--esp. in SV who make a point of dressing poorly and driving TESLA cybertroons--have come to attribute this phenomenon to superior to human taste, intelligent engineers tend to credit their superior dispositions, which as it stands, are yet to be baked into the weights via pre-training, or, circumvented entirely by apeing the human steers via RL maxxing. This is the sample efficiency problem which is rightly characterised as an all consuming event horizon around a proverbial black hole where data gets sublimated.

Let us consider an alternative to AI-HUMAN ensemble to better understand the premise which facilitates the type of judgement and taste that "become" multiplicative in terms of value added to agents and the work they do, which for this article, can be any kind of "ambitious" knowledge work (not merely laborious).

Alloyed agents

Here's a concept, pioneered by the good folks at XBOW.

In this experiment, XBOW tested performance on their own cyber-vulnerability benchmark, using, and get this, interleaved models from two different families (Sonnet from Anthropic and Gemini from Google) within the same loop, same chat thread, one assistant, the model(s) behind it alternating call to call, neither aware the other wrote some of its turns (ooo00 how hegalian!1).

From this article and experiment, I want you to hold two empirical observations in your head:

  • First, rather obvious are the gains from 25% to 55% solve rate, but these ONLY hold, when you conjoin two different model families rather than incestuously marry Sonnet 3.7 with Sonnet 4.0 (avg.)
  • The most de-correlated pair, Gemini and Sonnet, in terms of what each model found challenging by not solving, when interleaved, gave the largest uplift.

Complementary intelligences

So two models, from different families, serve as complimentary intelligences to each other. Gemini and Sonnet, within the same species of transformer based intelligence, will come with their own DISTINCT priors, and we can infer that these priors are substantial and obviously not trivial, because material differences arise in terms of task coverage.

The human prior

There is ofc, the perennial final boss of complimentary intelligence to AIs i.e. the human.

Humans, usually tend to differ from SOTA AIs, not only in terms of "normalised" parameter counts--Fable seeeeems to be 6 trillion, whilst you and yours truly sit at a comfortable lead of 100 trillion...plus), but also, thanks to evolution, orders of magnitude more sample efficient than AIs.

Here, is now where I'd like to articulate my points from earlier thesis.