The Thesis Failed. The Architecture Didn't.
What $500 of side-project R&D taught me that no vendor demo ever has
It took me weeks to build my thesis and one benchmark run to break it. I published the loss anyway.
This is a story about R&D, told from inside a profession that mostly doesn't do R&D. When I tell law firms and legal departments they need someone tinkering, someone trying things that might not work, I hear the same two objections every time: too expensive, and no time. So I want to walk through what one side project, built on nights and weekends for under five hundred dollars, actually produced. A failed thesis. A working governance architecture. A public benchmark result I would rather not have gotten. And more practical learning than any vendor demo has ever given me.
It up for anyone to view, Bailey. The failure is not the embarrassing part of that list. It's the productive part.
The tradition I was borrowing from
Legal AI is mostly sold to us in closed boxes. You get the demo, the pilot, the seat license, and a quality claim you cannot inspect. What you don't get is the machine itself.
There has always been a counter-tradition. In 2017, Mike Bommarito and Dan Katz took ContraxSuite, a contract analytics platform their company LexPredict had built, and open-sourced it, on the theory that software gets valuable when users can revise it. They did the same with LexNLP, a legal NLP library that outlived the company that made it. That was a strange thing to do in legal tech then. It still is.
A new crop is doing it anyway. Antti Innanen shipped Lavern, a multi-agent legal system, with a README that says the quiet part out loud: "It is at least ten things, several of which are, on their own, products somebody could build a company around. They are sitting in the repo. Take whichever ones you want." Will Chen is building Mike, an open-source legal AI platform, in public, right now.
That framing, ideas in a repo, yours to take, is the whole game. You don't start from zero. I didn't. Bailey runs on paperclip, a general-purpose agent control plane someone else built and open-sourced. I never forked it, never modified a line of it. I layered law on top: conflicts walls, citation gates, a deterministic deadline engine, receipts. The stack I'm describing exists because four or five groups of people published their work instead of hoarding it.
The bet
Here was my thesis, stated the way I'd have defended it at a dinner table last spring: legal work done by an orchestrated team of small, focused agents will beat the same work done by one strong model in one shot.
The mechanics felt obvious. A chat window that drafts a whole contract gives you one big opaque blob to trust or distrust. Small units seemed better on every axis I cared about as a former litigator: a focused agent doing one bounded task is easier to evaluate, easier to gate, easier to swap out when it underperforms. So I built the firm that thesis implied. A chief of staff that takes intake and delegates. Practice leads. About 180 agents and 178 skills in an org chart, from NDA triage to conflicts screening to renewal tracking, each one small on purpose.
The honest structure of that bet had two halves, and keeping them separate is the most important sentence in this essay. One half was about control: small units mean every consequential action can be classified, gated, logged, and receipted. That half is architecture. It's either built or it isn't. The other half was about quality: decomposed work will score better than monolithic work. That half is an empirical claim, and empirical claims deserve tests you can lose.
The test I could lose
Harvey, one of the biggest closed legal AI companies, publishes an open benchmark called LAB: over 1,700 real legal tasks with grading rubrics, plus their own automated grader. Which means a solo builder can score his homemade system with the same ruler a nine-figure lab uses. That sentence would have been science fiction five years ago.
I pre-registered the rule before running anything, because I know what motivated reasoning looks like in myself. If the orchestrated firm consistently beat the single agent, the thesis was supported. If they came out at parity, it wasn't. Then I ran the A/B: same tasks, same underlying models, one arm forced to work as a single agent, the other free to delegate through the whole org chart, everything scored by Harvey's grader, not mine.
The single agent never cleared ten percent of the rubric criteria. Nine judged runs across immigration, real estate, and banking tasks. Not once.
The orchestrated firm mostly sat at the same floor. Under the pre-registered rule, that's parity. Not supported at the tiers tested.
There was one genuinely interesting wrinkle. The only runs in the entire campaign that ever scored high, twenty-two of twenty-seven criteria on an immigration task, thirty-four of forty-six on a GDPR analysis, came from the orchestrated arm. Every single outlier. Orchestration showed a higher ceiling and an equal median: when the team got time and the delegation chain ran deep, it occasionally produced the best work anyone managed. It just couldn't do it reliably at the model tiers I could afford to run. And "could afford" is doing real work in that sentence. The whole campaign, judged runs included, cost less than a nice dinner. The entire project came in under five hundred dollars. That's the R&D budget that firms tell me is impossible.
The objection I owe an answer
If the benchmark failed, why should anyone touch this thing?
Because of what, precisely, failed. The claim that died was narrow: that decomposition produces better outputs than a strong single model, at the model tiers a $500 budget buys. Frontier tiers remain untested, the ceiling signal is real, and I published the harness so anyone with a bigger budget can run the next round. What never failed, because it was never on trial, is the governance layer. Gates, receipts, walls, and budgets are not hypotheses. They’re plumbing. And I’d argue the negative result is exactly why you can believe the rest of the repo: a system whose entire pitch is verifiable receipts should be able to show you the receipt for its own worst day.
The other honest objection is scale. Nine curated tasks, a handful of judged runs per cell, tiers bounded by my wallet. All true. This is a test case, not a verdict on orchestration. That’s what R&D produces: not certainty, direction.
What survived the test
Here's why I didn't delete the repo.
The quality half of the bet failed its test. The control half never depended on it. When my agents work a matter, every handoff is itself a matter, with an owner and a running token meter. Every agent runs under a hard budget cap it cannot exceed. When any agent tries to send something out of the building, a court filing, a signature request, a document upload, it stops at an egress gate and waits for a human, with the payload reduced to a hash. Every gate decision lands in a hash-chained receipt you can verify offline with openssl, no vendor in the loop, with optional trusted timestamping on top. Ethical walls run as separate tenants with separate gates and separate receipt chains, and a screened matter tells you it's screened instead of pretending not to exist. Filing deadlines come from a deterministic calculator that reports a blocker rather than guessing.
None of that is an aspiration. It's just built, and it works the same whether the agent behind it is brilliant or mediocre. Arguably it matters more when the agent is mediocre.
The test even changed the product's name. I almost called this thing Atomos, after the atomic-work thesis. But you don't name a project after the claim your own benchmark declined to support. A bailey is the walled courtyard of a castle, the place where work happens inside the walls, behind the gate. That's the part that held. So that's the name: Bailey.
One morning in August I typed a single request into the board: review this vendor's subscription agreement before signature, flag the renewal mechanics, the data rights, the liability posture. Then I did something that still feels strange in a chat-window world. I stopped typing, and I watched.
The chief of staff read the matter and opened two sub-matters underneath it. A risk memo went to the commercial lead. A second-pass review went to a risk spotter whose entire job is catching what the first pass missed. Neither of them asked me anything. Both left a trail, every handoff its own matter, with an owner, a status, and a meter.
Thirteen minutes and thirty-one seconds later, the parent matter turned green. The consolidated deliverable opened with the conflicts notice the policy layer refuses to skip, then got to the point: Delaware law, twelve-month initial term, auto-renewing, recommended posture, do not sign as-is. About 1.8 million tokens of work. Pennies.
Then one of the agents tried to leave the building. A court-filing request hit the egress gate and stopped cold. My approvals board showed the request, the payload represented only by its SHA-256 hash, and two buttons: approve, reject. Nothing was going anywhere until a human clicked.
And at the bottom of all of it sat the receipt chain, with a head hash I verified on my own laptop. While I was reviewing it, I noticed the system had also, entirely on its own initiative, requested permission to hire two more agents it thought the firm needed. Requested. Permission. It's still waiting.
That half hour taught me more about what supervising AI work actually feels like than two years of product demos.
Take it
Bailey is on GitHub under Apache 2.0, published the way Antti published Lavern: as a source of inspiration. It's a test case. I may keep developing it. I may not. It makes no representations about any end use, and anything touching real legal work needs a lawyer's judgment wrapped around it. Take the entire thing, or take the parts, the gate proxy, the receipt spec, the walls, the deadline engine, and make them your own. Hand something forward when you're done.
But the repo is almost beside the point. The point is that the two objections, too expensive and no time, did not survive contact with a side project. Under five hundred dollars bought a full experiment cycle: a real bet, a real build, a real test against a grader I don't control, a real failure, and a real discovery about what actually holds up. Iteration in the open isn't the confession you make when things go wrong. It's the method. You keep going, and going, until you figure out something that's helpful, and then you show your work, including the parts that broke.


