Economics 9 min read

The economics of the harness — why we run frontier models

There’s a story running through the industry press right now, and it goes like this: frontier AI models are astonishing but ruinously expensive, so serious companies are migrating their workloads to cheaper models and accepting whatever intelligence they lose in the trade. The numbers in those stories are real. Teams running open-ended AI workloads at scale really do see bills that force the question.

We read those stories with interest, because our experience is the opposite — and not because we use the AI sparingly. Scout runs on frontier models, the most capable readers money can rent, and model cost has never been the pressure that shaped the product. That isn’t luck, and it isn’t scale we haven’t hit yet. It’s a consequence of the same architecture this site keeps describing for other reasons: the model reads; code decides. It turns out that the rule that makes a review reproducible is also the rule that makes it affordable.

This article is the accounting.

The bill is mostly the words coming back

Model providers price two things separately: the tokens you send in — the documents, the instructions, the questions — and the tokens the model generates back. Output is priced several times higher than input, because generating is the expensive part. So the economics of any AI product come down to one question: how much does your system make the model say?

A chatbot’s answer is: a lot. Its entire product is the model talking — paragraphs of prose, drafted and redrafted, every word billed at the premium rate. An “agent” that reasons out loud, explains itself, writes and rewrites its plan? More talking still. When those workloads get expensive, there’s nothing to trim except the model itself. That’s the corner the downgrade stories are written from.

Scout’s answer is: almost nothing. The harness asks the model to do one job — read pages and hand back specific facts. The request going in is big: real scanned documents, whole pages of them. The answer coming back is a handful of terse, structured values — an APR from a disclosure, the guarantors on a commercial note, the transaction facts that decide whether a currency transaction report is owed. No essays. No explanations. No reasoning performed for an audience.

And everything downstream of the reading is deterministic code, which produces its output for free, the way software always has. That’s true well beyond the rule check at the end. A recipe orchestrates — it sequences the steps, calls the tools, moves facts between systems, assembles the finished report — and orchestration is program logic, not conversation. As recipes have grown past QC into commercial file review and deposit-operations work, the work per run has grown; the model’s share of it hasn’t. Whether the recipe is checking an indirect auto file, walking a commercial loan’s covenant documents, or screening the day’s large cash transactions, the model is consulted the same way — read this, return these facts — and everything around those calls is code.

The shape of our spend reflects that: month over month, for every twenty-plus tokens Scout sends to the model, roughly one comes back. The expensive direction is the one the architecture barely uses. We didn’t achieve that ratio by economizing. We achieved it by never asking the model to do anything except the thing it’s uniquely good at.

Structure pays on the input side too. Model providers bill previously-seen input at a steep discount — prompt caching, in the trade — but only when requests repeat exactly, and open-ended conversations rarely do. Scout’s requests repeat by construction: the instruction is a constant in the code, and a recipe asks its questions the same way on the thousandth loan as on the first. The discipline that exists so reviews are reproducible turns out to be exactly what the discount was priced for. Sprawl can’t cache; structure can’t help but cache.

A recipe is a token budget you can read

Structure doesn’t just make the spend small. It makes it predictable — and for a bank negotiating a contract, predictable matters more.

A review in Scout is run by a recipe — a sealed, versioned program with a fixed extraction checklist and fixed rules. That means a given review has a shape: this many document types, this many facts, asked for in attention-sized batches, corroborated, done. Run it on a hundred loans and you get a hundred similarly-sized token bills, because the same program did the same work a hundred times. And the unit prices independently per recipe: the consumer QC check, the commercial file review, the deposit-ops screen each have their own profile, so a new use of Scout arrives with a knowable cost instead of widening an unknowable one. The ledger makes this visible rather than theoretical — every model call is recorded with its token counts, and no call site can escape the meter — so “what does a review cost?” is answered from records, not estimated from vibes.

Compare that with the open-ended alternative, where cost is a function of how long the conversation wandered, how many times the agent re-read its own reasoning, how verbose the model felt. Nobody can price that per unit, which is why so much AI pricing is metered anxiety.

A sealed recipe doesn’t just produce the same findings every run. It produces roughly the same bill — which is the difference between a cost and a price.

This is what notebooks and recipes were built for — repeatability of process — and pricing predictability falls out of it as a side effect. When the process is a program, its cost profile is a property of the program, not of the weather.

Which is why we upgrade models instead of downgrading them

Here’s the part of the downgrade story that deserves more scrutiny than it gets: when a team swaps a frontier model for a cheaper one, the thing they’re economizing on is reading comprehension. In a compliance product, that’s the one line item where cheap is expensive. A misread deductible, a hallucinated VIN, a guarantor dropped from a commercial note, a disclosure field silently skipped — the harness catches these (that’s its job), but every catch is a retry, an escalation, or a human interruption. Weaker reading doesn’t make a review cheaper. It makes it slower and noisier.

So Scout’s posture runs the other way. Because structure — not model restraint — controls the spend, upgrading to a newer, smarter model is something we get to do eagerly. The harness already runs a ladder: fast, inexpensive models take the first pass at routine classification, and a stronger model is brought in when the first pass comes up empty. When a provider ships a better frontier model, adopting it is a configuration change inside machinery built to be model-agnostic — the prompts, the batch sizes, the verification gauntlet, the ledger all stay put. The reading gets sharper; the economics barely move, because the model was never being asked to talk.

We like frontier intelligence. We think banks deserve it — the documents are messy, the stakes are regulatory, and “almost as good at reading” is not a savings. The harness is what makes that appetite affordable: it buys the best reader available and then refuses to let it do anything expensive.

Choosing whose model to rent

One more economic decision hides inside “the model is rented capability”: rented from whom? A model provider is a vendor — the same category of decision as a core processor or an item-processing outsourcer — and it deserves the same third-party diligence, because your examiner will treat it that way. The questions are the standard ones. Where does the processing physically run? Under whose law does the provider operate, and whose courts enforce the contract? What does the retention agreement actually say? Who is accountable — findable, contractable, auditable — when something goes wrong?

Scout’s answers are already part of the shipped promise: inference contracted to run in the United States, under a zero-data-retention agreement, with nothing used for training — from providers whose contracts are enforceable where our customers are examined. Those weren’t the only options, and they weren’t always the cheapest. They were the ones whose answers we could put in writing in front of a regulator.

Zero data retention deserves one more sentence, because it’s usually read as a privacy checkbox when it’s really a risk posture. It means that when a request completes, the provider holds nothing — the pages were read in the moment and never stored. So the questions that follow every stored copy of customer data — who can access it, how long it’s kept, what happens if that store is breached, who else’s subpoena can reach it — don’t get careful answers here. They get no object. There is no store to breach, no archive to subpoena, no dataset to leak. Nobody can lose what nobody possesses. That’s a stronger position than any promise about how well a copy is protected, and it applies to a commercial borrower’s tax returns and a depositor’s account records exactly as it applies to a consumer loan file — the boundary doesn’t know or care which line of business the page came from.

That’s the whole position. Not a judgment about any model or any country — capable models are being built in many places, and the field is better for it. It’s the observation that institutions that wouldn’t move item processing offshore without a long conversation shouldn’t move document comprehension offshore without the same one. We chose vendors so that the conversation is short.

The axis that actually matters

The industry is currently arguing about cost versus intelligence, as if those were the ends of the dial. They aren’t. The dial is structure versus sprawl. A sprawling system — open-ended conversations, reasoning performed at retail prices, output nobody budgeted — is expensive at any model tier, and downgrading the model just makes it cheap and worse. A structured system is economical at every tier, which frees it to run the best one.

Scout spends where reading quality lives, saves where determinism is better anyway, and can tell you the price of a review before you run it — with a ledger to prove it after. That’s not a cost-control program bolted onto an AI product. It’s what the harness was for all along.