deterministic labs
← all writing

Writing, August 2026

Ali Khalilli

Firm ground

Why our machines sound right when they are wrong, and what to do about it.

A clock that runs exactly one hour fast is never right. Not once a day, not once a year; it is always one hour ahead of the truth. A stopped clock, at least, is right twice a day.

Both clocks are perfectly consistent, each does the same thing every time we look. But: one is sometimes right and one never is. So consistency and correctness are different things. A machine can be perfectly consistent and perfectly wrong at the same time.

If we look at which clock is the dangerous one? The stopped clock fools nobody: anyone can see it is dead. The running clock ticks, spins, and looks alive. And it is wrong every single time.

Fig. 1Two consistent clocks

Stopped

right twice a day

One hour fast

never right

Each shows the same time whenever we look. Both are perfectly consistent; only one is sometimes correct, and it is the one that looks dead.

AI today is becoming that running clock. The trick is old, the scale is new. And there is a reason it works on us. Psychologists describe two modes of human judgment: a fast one that runs on impressions, and a slow one that checks[1]1Kahneman, Thinking, Fast and Slow (2011). Two modes of judgment: one fast, one that checks.. The fast mode has a favorite piece of evidence, also we can say it’s single signal: fluency. What reads smoothly feels true[2]2Alter & Oppenheimer (2009). What reads smoothly is felt as true.. A confident, wrong answer from a machine does not only fail to warn us about that. It arrives in exactly the format our fast judgment(system-1) accepts without inspection, and the slow judgment, the one which checks, is never called. The machine is not lying. It does that without intending anything, speaking straight to the part of us that does not check. The smoother the machine becomes, the harder the mistake is to see.

Our name

We are called Deterministic Labs. Deterministic /dɪˌtɜː.mɪˈnɪs.tɪk/[3]3“Deterministic,” Cambridge Dictionary. is a word with a simple meaning: same input, same output, every time. A fingerprint function like SHA-256 is deterministic. Feed it a file and it returns the same sixty-four characters, today, tomorrow, on any machine on Earth. Change a single comma in the file, and the fingerprint transforms beyond recognition. That is determinism: a promise so strict that the whole world can hold you to it.

Here is the “surprise” in our name: we do not believe everything should work that way.

Determinism is valuable the way a foundation is valuable. A foundation is a wonderful thing, but you do not pour concrete everywhere. You pour it where the ground is solid. The first skill in this field is telling solid ground from mud.

Our name also has a second meaning. We will explain it at the end.

One answer, several layers

Take the number TWELVE. We can write it as 12, as twelve, as a dozen, as XII. Four ways of writing it. One number.

Suppose an AI model answers a question with “twelve” today and “a dozen” tomorrow. Did the answer change? No. The number stayed the same. Only the wording changed, and the wording was allowed to change.

So “the same answer every time” is not one requirement. It is several requirements, stacked:

Fig. 2One answer, five layers (decide which layers must not change, lock them, and let the others stay free)
question

How many months are in a year?

answer
There are twelve months in a year.
what you can read ↑
tap a layer to lock or free it
↓ what the machine does
run 07, seed 07
Twelve, read at five heights, drawn as an iceberg: solid is locked, open water is free. Ask again and the free layers vary while the locked layers repeat; every run leaves a receipt. Tap a layer to lock or free it.

The same answer can be locked on some layers and free on others, at once. That gives the whole craft in one sentence: decide which layers must not change, lock them, and let the others stay free.

For the number twelve, lock the meaning and free the wording. An AI forbidden to say “a dozen” is not more precise. It is worse at language. The reverse works too: a report can have lively writing and exact numbers at the same time. The style is free. The math is locked. Different rules on different layers, inside one answer.

Most arguments about whether AI is “reliable enough” are really two people talking about two different layers without noticing.

Where the randomness enters

People say AI model is random. The truth is more precise, and more useful.

The model itself is fixed. Give it the same question twice and it produces the same internal result twice: a list of possible next words, each with a probability. Same list, same numbers, every time. The randomness enters in one extra step at the end. The system picks from that list partly by chance, like a weighted roll of dice. The chance is there on purpose: always choosing the top word makes the writing flat and repetitive[4]4Holtzman et al., ICLR 2020. Always taking the top word degrades the text.. That roll is the CHOICE layer. Remove the dice, always take the top option, and almost all the randomness disappears.

Almost all. A tiny amount remains, down at the MATTER layer, and its cause is surprisingly ordinary. Computers add numbers in a specific order, and changing the order can change the last decimal places[5]5Goldberg (1991). Floating-point addition is not associative.. When many people use an AI service at once, their requests are processed in groups, and the size of the group changes the order of the arithmetic[6]6Thinking Machines Lab (2025). Batch-invariant kernels pin the order and remove the randomness/noise.. So your answer can change depending on how many strangers pressed enter at the same moment as you. Even this can be fixed. Engineers can force the arithmetic into one fixed order, whatever the group size[6]. It costs a little speed, and on any one machine, set up one way, the wobble is gone entirely.

Fig. 3One question, asked again (the model never changes; find what does)
The question
every output is a ________.
The model: fixed
the model
asked ×0, changed ×0
The list: matter
Choice: the dice
a weighted roll, on purpose
The answer
________
not asked yet
The model computes the same list on every prompt. The spread we see is the dice, rolled on purpose. But a second randomness hides in the last decimals: other people’s requests, running at the same time, change the order the sums are added in. Take the top word and pin the order, and the same question returns the same answer, every time.

We tell this small story because it is the whole domain in one small example. Exactness is not something computers give us for free. It is something we build, after we understand precisely what was moving and why. “Computers are just random sometimes” is a give up & excuse. “This sum, in this order, pinned down” is firm ground.

One clarification, before the bigger idea. Deterministic does not mean predictable. The weather follows fixed physical laws, and still nobody can forecast it three weeks ahead, because tiny differences today grow into huge differences later[7]7Lorenz (1963). Fixed laws, yet unforecastable: the weather.. A system can follow strict rules and still surprise everyone.

Three kinds of questions

Now the bigger idea, and the heart of this essay. Questions come in three kinds, and each kind can carry a different amount of certainty. The worst mistakes in AI come from treating all questions as one kind.

Kind one: solid ground. Is 2,147,483,647 a prime number? It is. Euler proved it by hand in 1772[8]8Euler, 1772 (rec. Dickson, 1919). 2,147,483,647 is prime., and programmers meet it daily as the largest number a 32-bit computer can hold. No opinion, no style, no confidence can make it otherwise. Questions of this kind have one right answer, checkable against something outside the AI model: a calculation, a proof. Here, demand exactness. Build in stone.

Kind two: solid ground with a hidden condition. Ask “what temperature does water boil at?” and the honest answer is a question back: at what altitude? At sea level, 100 degrees. On a mountain, less. The answer was always exact. It only needed one condition named first. “Is this a good salary?” has the same shape: name the city and the profession, and a vague question becomes a lookup. A great deal of what looks vague works exactly this way: a crisp answer, waiting for a condition that no one stated. Name the condition, and the ground turns solid.

Kind three: soft ground. Is this person tall? Is this poem beautiful? There is no answer key for these anywhere, and there never will be, because “tall” is not a fact about the person. It is a comparison to a standard that you chose. Tall next to whom? Change the standard and the answer changes, and that is correct. Demanding one exact answer here is not rigor. It is a misunderstanding of the question.

Fig. 4Three kinds of ground (press a sky to drop a question)
press the sky to ask
build in stonename the condition firstan honest opinion, labeled
press the sky to ask

is 2,147,483,647 prime?

at what temperature does water boil?

the fog asks back: at what altitude?

hidden condition

is this person tall?

One sky, three grounds. On stone every drop lands on the same exact line. In fog the stone dissolves until you name the hidden condition; drag the altitude and the answer turns exact. On the dune each drop settles differently, an honest opinion against the standard you chose. The guarantee must match the ground.

You would not trust a calculator that got creative. You would not want a poet who wrote the same poem every time. The guarantee must match the ground.

And this is the failure we are here to fight: treating soft ground as solid. Pouring concrete over mud, stamping it “verified,” and selling it with the confidence of the fast clock. The opposite failure is real too: treating solid ground as soft, and answering a question of plain fact with a vague opinion. Both failures come from the same blindness. Not knowing, or not caring, what kind of ground you stand on.

The line moves

One more fact, and it is the most important one here. The line between “checkable” and “matter of opinion” is not fixed. People move it, deliberately, by building tools.

Whether a mathematical proof was correct used to be an expert’s opinion. Reviewers read it and judged, and reviewers disagreed. The four-color theorem was contested for years because its original 1976 proof used a computer that no human could fully audit. Then proofs were given a form that programs can verify, and a machine now checks the entire theorem, every line, in minutes[9]9Gonthier (2008). The four-color theorem, checked by machine.. A check nobody trusted became a check anybody can run. The question moved from opinion to fact. Whether software is correct is making the same journey right now, piece by piece, wherever someone pays the cost of writing down exactly what “correct” means[10]10Leroy (2009). A C compiler proven correct against its specification.. In every case, the mud did not dry on its own. Someone built solid ground under it.

This is the second meaning of our name. Deterministic is not only what we look for. It is what we do: we make things checkable that were not. There are many questions in AI today that receive a confident opinion when they could receive a check. Our work is moving them, one by one, from “trust me” to “check it.”

What we know, and what we bet

An essay that asks machines to label their claims should label its own.

What we know: consistency and correctness are different things. “The same answer” has layers, and each layer can be locked separately. The randomness in AI comes almost entirely from one sampling step, plus a small hardware quirk, and both can be controlled. And questions come in three kinds, each bearing a different kind of certainty.

What we bet: the industry is perfecting the sound of intelligence, and the sound is nearly perfect. Our first bet is about skill. The scarce ability of the coming years is not fluency but discrimination: knowing what kind of question is on the table, and answering with the kind of certainty that question can carry. Our second bet is about what follows from that skill. Trust will reorganize this market. Systems that can be checked will be given decisions. Systems that must be believed will be given suggestions. The distance between those two roles is where the value of machine intelligence will settle.

We may be wrong. That is what makes them bets.

Where we stand

Today’s AI answers a calculation, an estimate, and a matter of taste in one identical, confident voice, and gives you no way to tell which is which. That is the fast clock at full scale: a machine that is wrong in the same tone it uses when it is right. And the tone is not an accident. These systems are tuned on human approval, and human approval follows the fast judgment: we reward what sounds right[11]11Sharma et al., ICLR 2024. Tuned on approval, models learn to sound right.. Tune a machine long enough on the sound of rightness, and you get a machine optimized for the sound. Being right survives the tuning mostly where it is also easy to check.

Its answers are made to be believed. Ours are made to be checked. Everything we release will arrive with the means to check it. Run it yourself. Measure it yourself. Compare it yourself. We begin at the bottom. From there we build upward, layer by layer, and each layer will stand on a checked one.

For us, determinism means three things. Knowing where certainty is real. Being honest about where it is not. And moving the line between them.

It is the difference between the clock that is confidently never right, and the one that knows its limits and keeps time.

Everything we make begins here.

References

  1. Kahneman, D. Thinking, Fast and Slow. Farrar, Straus and Giroux, 2011.
  2. Alter, A. L., and Oppenheimer, D. M. “Uniting the Tribes of Fluency to Form a Metacognitive Nation.” Personality and Social Psychology Review, 13(3), 219–235, 2009. doi.org/10.1177/1088868309341564
  3. “Deterministic.” Cambridge Dictionary. dictionary.cambridge.org/dictionary/english/deterministic
  4. Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. “The Curious Case of Neural Text Degeneration.” ICLR, 2020. arxiv.org/abs/1904.09751
  5. Goldberg, D. “What Every Computer Scientist Should Know About Floating-Point Arithmetic.” ACM Computing Surveys, 23(1), 5–48, 1991. doi.org/10.1145/103162.103163
  6. He, Horace, and Thinking Machines Lab. “Defeating Nondeterminism in LLM Inference.” Thinking Machines Lab: Connectionism, September 2025. thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference
  7. Lorenz, E. N. “Deterministic Nonperiodic Flow.” Journal of the Atmospheric Sciences, 20(2), 130–141, 1963. journals.ametsoc.org/view/journals/atsc/20/2/1520-0469_1963_020_0130_dnf_2_0_co_2.xml
  8. Dickson, L. E. History of the Theory of Numbers, Volume I: Divisibility and Primality. Carnegie Institution of Washington, 1919. Records Euler’s 1772 proof of the primality of 2³¹ − 1. archive.org/details/historyoftheoryo01dick
  9. Gonthier, G. “Formal Proof: The Four-Color Theorem.” Notices of the American Mathematical Society, 55(11), 1382–1393, 2008. ams.org/notices/200811/tx081101382p.pdf
  10. Leroy, X. “Formal Verification of a Realistic Compiler.” Communications of the ACM, 52(7), 107–115, 2009. xavierleroy.org/publi/compcert-CACM.pdf
  11. Sharma, M., et al. “Towards Understanding Sycophancy in Language Models.” ICLR, 2024. arxiv.org/abs/2310.13548