bio | photo stream | social | email

technology | photography | games | other

← Back to all posts

Cognition vs Determinism

Musing on agent design patterns for success, safety, and cost

September 16, 2026

We've come a long way since the early days of LLM powered chatbots, yet you still read stories of bots handing out wild discount codes and cancelling other peoples orders. These are basic mistakes with pretty clear solutions.

As models have become "smarter" with the way they operate with reasoning, frontier providers have removed knobs like temperature and top_p that let you control determinism. This is generally fine except if you need mission critical operations, like a user asking a chatbot to cancel their order or update their credit card. If a bot cancels the wrong order, or instead responds with the users credit card number ... or another users credit card number, big problems happen. It is wildly inappropriate to build such an agent.

So the solution when designing an agent is to build tools. A tool that gets info on a user. A tool that gets a list of users orders and their status. A tool that cancels an order. A tool that securely updates a credit card number. Those tools are called by the LLM, but with information provided deterministically by the agent harness you are building. An LLM doesn't provide the user ID into those tools, it is provided by the software that authenticated that user. For things like order numbers you would directly pass an order ID from a "find order" tool to a "cancel order" tool. Don't let the agent fat finger it. Build safety into both tools that check against the user ID provided by the authentication that is performed before the LLM even replies. Now an LLM can't accidentally query the wrong user, or cancel the wrong order. By reducing the cognition to simply be tool routing you've greatly improved the determinism of your solution.

Because of this you can now likely run your agent on a cheaper / faster model too, and with less reasoning. It simply needs to choose which tools to summon to meet the users request, not to actually try to solve anything. Reasoning may in fact be your enemy here, so benchmark with as little as possible. Plenty of benchmarks measure the relative tool call reliability across models, and plenty of very affordable models dominate those charts. You don't need Fable, Opus, maybe even Sonnet to be your LLM in this scenario. Models like Gemini Flash are excellent tool routers, simply providing the human interface to software (written by a smarter LLM).

If your cost and latency budget allows you can deploy a validator pass on LLM replies before they're sent on to the user. This should be another LLM, potentially another model, performing rapid QA against the reply. If it fails the validator (i.e. we went off script tonally) you can retry the original request however many times is appropriate, ultimately erroring instead of replying incorrectly to a user.

Something I had talked about in a previous little blog post was how critical logging is. Write every call to bigquery, the context, the tool calls made, reasoning. Figure out how long you can cost effectively retain this. You can both use this for QA and agent improvement, but also model regression testing when you decide to move the agent models need to change.

Our team is having an absolute blast writing agents, hopefully you are too!

← Back to all posts