Note: The Listed platform will permanently shut down December 31, 2026. Your data published on Listed will still be available in your personal Standard Notes account. Learn more

AI Coding Agents Need a Driving Test, Not a Beauty Pageant

Databricks just published one of the more useful looks at AI coding agents because it asks the question companies should actually care about: not “which model sounds smartest,” but “which one can fix our code without lighting money on fire?”

The post, Benchmarking Coding Agents on Databricks’ Multi-Million Line Codebase, is worth reading because it moves the discussion away from leaderboard theater and toward real engineering work.

That distinction matters.

Public benchmarks are useful, but they are not your codebase. They do not know your weird service boundaries, your old dependencies, your new experiments, your sacred config files or the one module nobody touches unless absolutely necessary. Databricks tested coding agents against its own multi-million-line codebase, using real engineering tasks from actual pull requests. That is a much better test than asking an AI to solve some public repo problem it may have indirectly seen before.

In my experience, this is exactly where the rubber meets the road. A coding agent can look brilliant on a clean demo and then become a confused intern the second it hits a real project with hardware quirks, stale comments, half-finished ideas, renamed services, old assumptions and three different ways of doing the same thing. The model matters, but the workflow matters just as much. How much context does it get? Can it see the right files? Does it understand the actual architecture, or is it just confidently patching whatever file happened to be open?

I ran into some of these problems using agents myself. The catch is that I cannot really run those agents through my normal chat subscriptions. Once you get into agentic coding, tool calls and API usage, you are burning tokens. The meter is running.

I started using OpenClaw on a Raspberry Pi 4B, and that is where the “cheap” model started feeling expensive in a different way. It kept asking me to run tasks, paste results back, run another command, report back again and then wait while it guessed what the output meant. At some point I thought: I am paying for this thing, and now I work for it.

The more expensive model was much better. That is the uncomfortable part. The cheap model was cheaper until it was not. I was spending about 80 cents a day and started watching the number. Some days it was $1.50. That was enough to make me pay attention. Then I did a bunch of work with it one day and hit $5. That was not catastrophic, but it was enough to make the warning lights come on.

Then came the real problem: it had decided to run every heartbeat. It was not doing much useful work, but it was still burning tokens. That is the kind of failure that does not look dramatic in a demo. Nobody screams. Nothing explodes. The agent just quietly turns into a tiny cloud-based furnace for your money.

That is where “cheap” gets complicated.

The results from Databricks are less “one model to rule them all” and more “use the right wrench.” Databricks found that top performance came from a mix of OpenAI, Anthropic and open-source models. That is the first useful lesson. The future of AI-assisted engineering is probably not one magic chatbot sitting on a throne. It is routing. Easy task? Use the cheaper model. Ugly architectural goblin problem? Bring out the expensive brain.

The second useful lesson is that open models are no longer just the guy at the party saying Linux is better while standing next to the snack table. Databricks found GLM 5.2 performed in its top capability tier and was competitive with high-end proprietary models in its benchmark. That does not mean every company should immediately switch everything to GLM. It does mean the old assumption that open models are always second-class coding assistants is getting stale fast.

That is one reason I switched to Hermes. I am running it on the Pi and on a VPS, and it has a free model that is capable enough to be useful. That changes the economics. It does not mean every task should go to the free model, but it gives me another gear. Some jobs do not need the expensive brain. Some jobs just need something competent enough to read the room and not set the token meter on fire.

Token price is not task price

A model can be cheaper per token and still more expensive per job if it burns through context, loops too long, reads half the repository and then confidently edits the wrong file. I have seen this kind of thing in smaller projects too. Sometimes the “cheap” answer is cheap because it gives you three more problems to debug afterward. Sometimes the expensive model is worth it because it gets the shape of the system right on the first pass.

There is another wrinkle for individual users and small operators: API pricing is not always the same calculation as using a chat-based subscription. I pay for ChatGPT and Claude, and sometimes the cheapest practical way for me to get the work done is not metered token API usage. It is the monthly subscription I am already paying for. Claude has weekly limits and shorter usage windows, and I can run into that structure differently than I do with ChatGPT. With ChatGPT, I have not hit those limits in a while.

That matters because “cost” is not just a spreadsheet number. It is money, time, friction, waiting, context switching and whether the tool is helping me move faster or making me feed it breadcrumbs like a confused raccoon in a server room.

Databricks found the same general problem at larger scale. Token price by itself was a poor indicator of actual cost. A model may look cheaper on paper and still cost more per completed task if it works longer, reads more context and burns more tokens to get there.

That is the AI equivalent of buying the cheap printer and discovering the ink cartridges require a home equity loan.

The harness matters

The fourth lesson is that the agent harness matters as much as the model. The wrapper around the model, the tool loop, the amount of context it sends each turn and how tightly it manages the working set can change both cost and performance.

Databricks found that the same model could have dramatically different costs depending on the harness. That matches what I have seen: the best results usually come when the agent is boxed into a clear job, given the right files and forced to work against reality instead of vibes.

The best part of the Databricks post is the benchmarking method. They used real pull requests, removed the original solution, kept relevant tests aside, gave the agent a task prompt and then judged the result with tests rather than vibes. They also discovered a great little problem: because the tasks came from merged commits, some agents could find the real answer in git history. So they sealed off git history.

That is both hilarious and important.

The robot did not “solve” the problem. It found the answer key in the teacher’s desk.

Benchmark the work you actually do

The larger point is simple: AI coding agents are becoming real engineering tools, but they need real engineering evaluation. Companies should stop asking which model is “best” in the abstract.

Best for what? Cheapest by what measure? In which codebase? With what harness? On what kind of task? Judged by tests, humans or vibes in a blazer?

Databricks’ answer is the correct one: benchmark against the code you actually ship.

That is where the useful truth lives. Not in leaderboard theater. Not in AI demo magic. Not in a model confidently explaining a bug it invented three seconds ago.

Run the agents on your own work. Measure task completion, cost, context use and failure modes. Then route the simple jobs to the cheap fast tools and save the expensive models for the problems that deserve them.

In other words, stop treating AI coding agents like celebrity geniuses.

Treat them like contractors.

Give them a real job, hide the answer key and see who passes inspection.

And if the contractor makes you do all the work while billing you every heartbeat, fire the contractor.

More from dickie
All posts