Independent AI intelligence Two editions daily · ET
Fervor AI

Analysis · October 7, 2026 · concept

OpenAI Decisions APIgpt-6-lunaagent-infrastructureagent-securityfrontier-modelsagent-harness

OpenAI's Decisions API Bills Only What It Reads, and That Changes How You Write Agent Guards

No output charges, a higher input rate than the same model elsewhere, and why the context you send a guard is now both its price and its accuracy

OpenAI's newest API endpoint never writes a sentence. You hand it some text or an image and a question with fixed answers, and it hands back a probability. On October 6 it went to public beta with a pricing line that looks almost too simple: input costs $0.10 per million tokens, and output, cache reads and cache writes cost nothing.

Nothing for output sounds like a discount. Read the rest of OpenAI's own price list and it turns out to be something stranger, a bill that only cares about what the model reads. That one detail should change how you design the small checks that sit inside every agent loop.

What the Decisions API actually is

The Decisions API lives at POST /v1/decisions and runs on one model, gpt-6-luna. It answers three kinds of question:

  • A predicate returns the probability that a condition is true.
  • A choice picks one option from a list you supply and returns a probability for each option plus a confidence.
  • A score rates the input against ordered levels you define and returns a probability-weighted average.

OpenAI pitches it as about ten times faster than sending the same question through the Responses API. The answer comes back as typed fields (probability, choice, probabilities, confidence, score), so there is no parsing step and no chance the model wanders off into an explanation you did not ask for.

Here is the shape of a choice question, straight from the docs:

decision = client.decisions.create(
    model="gpt-6-luna",
    input="I was charged twice for my order.",
    questions=[{
        "type": "choice",
        "name": "department",
        "instructions": "Which department should handle this complaint?",
        "choices": [
            {"value": "billing", "description": "Payments, invoices, and refunds."},
            {"value": "technical", "description": "Problems using the product."},
            {"value": "shipping", "description": "Delivery and tracking."},
            {"value": "other", "description": "Requests outside these categories."},
        ],
    }],
)

If you have built an agent, you already know where this goes. Routing a task to the right subagent is a choice question. Deciding whether a shell command is safe to run is a predicate. Ranking a retrieved chunk by how well it answers the query is a score. These are the calls your agent makes most often, and until now they were usually tiny generation requests dressed up as classification.

The price shape is the interesting part

Here is the detail that took me a second read. OpenAI's pricing page lists gpt-6-luna through the regular API at $0.05 per million input tokens and $0.25 per million output tokens. The Decisions API charges $0.10 per million input tokens for the same model name and nothing for output.

So the input rate doubled and the output rate went to zero.

That trade has a break-even point you can work out on a napkin. Through the regular API, a call costs 0.05 × input + 0.25 × output (per million tokens). Through Decisions, it costs 0.10 × input. The two are equal when output is one fifth of input. If your current guard emits more than 20% as many tokens as it reads, which is easy once a model reasons out loud before answering, Decisions is cheaper. If your guard reads a long transcript and answers "yes," the regular API may well be cheaper per call, though you give up the speed and the typed probabilities.

The docs add one more wrinkle: regional processing premiums and long-context input multipliers apply on top. The page does not say how large those multipliers are. Long-context input is exactly where a careless guard ends up.

Put those together and the conclusion is plain. On this endpoint, the only thing you pay for is what you send. Every token of context you attach to a guard is a token you are billed for at double the regular rate, and nothing the model does after that costs you anything.

Why context is also the accuracy problem

If cost were the whole story, this would be a footnote about trimming prompts. The same lever also controls whether the guard is right.

A decision model returns a probability spread across the options you supplied. That is useful, and it is also limited in a specific way: it can only weigh what you put in front of it against the answers you wrote. Give a safety predicate the entire session transcript, and the question "does this command write outside the repository?" now competes with forty turns of unrelated chatter, earlier commands that were fine, and the user's own instructions. Give it the command, the working directory and the repository root, and the question has almost nowhere to hide.

OpenAI's guidance points the same way. The docs tell you to "use labeled examples from your application to set thresholds for routing, filtering, or review" and to "choose thresholds based on the cost of false positives and false negatives." That advice only works if the input to the guard is stable enough that a threshold set on Monday still means something on Friday. A guard fed an ever-growing transcript sees a different distribution every time the conversation gets longer.

Small open decision models show the same design. The Strands Decider 2B post, published by the Strands Agents team on October 1, describes the model sitting in an intervention handler that checks a tool call before it executes. Its authors say plainly that it is "significantly worse at solving complex problems than reasoning models." That is the right expectation for any judge in this class. It is fast because it does one narrow thing.

Put this into practice

Here is what I would do this week, in the order that costs the least effort.

1. Find your most frequent small call. Look at your agent logs for the request that fires most often and returns the least text. A tool-approval check, a router, a relevance filter. That is your first candidate.

2. Rewrite it as a question with fixed answers. If it is a yes/no, make it a predicate with one clear condition. If it picks a destination, make it a choice and give every option a one-line description, the way OpenAI's example does. Always include an "other" option so the model has somewhere honest to put inputs that fit nothing.

3. Cut the context to what the question needs. For a command guard, that is usually the command, the working directory and the allowed paths. For a router, the latest user message and the list of agents. Write down why each field is there. If you cannot say, drop it.

4. Measure both prices on your real traffic. Take a day of logged calls, count input and output tokens, and run the napkin math above. If your output is under a fifth of your input, the regular API may be cheaper and you should decide whether speed and typed probabilities are worth the difference.

5. Label fifty cases and set the threshold. Pull fifty real inputs, mark the right answer by hand, and pick the cutoff that matches what a false positive and a false negative cost you. Then log every score in production and audit the misses once a week.

6. Mind the image rule. If your guard looks at screenshots, the docs say images "must be inline base64 data URLs." Hosted URLs and file_id inputs are not supported, so plan the encoding step before you wire it in.

Honest limitations

This is a public beta with one model. If gpt-6-luna is wrong about your domain, there is no bigger sibling on this endpoint to fall back to, and OpenAI says general availability is only expected "in the coming weeks."

The pricing comparison above uses OpenAI's list prices as of this week. The Decisions page does not publish the size of its regional or long-context multipliers, so a guard that crosses the long-context line will cost more than the napkin math says, by an amount I could not pin down.

I have not run a calibration study of my own on gpt-6-luna. The ten-times speed claim is OpenAI's, and the threshold advice is OpenAI's. Treat both as starting points, not results.

A decision endpoint also cannot tell you that none of your options were right. A probability across four departments says nothing about whether the complaint belonged to a fifth. The "other" option helps. It does not fix a question that was badly framed in the first place.

And moving checks to a cheap endpoint invites you to add more of them. More guards means more places for a confident wrong answer to block good work or wave bad work through. Cheap calls multiply, and so do their errors.

Where this leaves your agent

The Decisions API puts a number on something builders have been fuzzy about: the small judgments inside an agent loop are a different workload from the reasoning, and they deserve their own design. The price shape makes the lesson hard to miss. You pay for what the guard reads and nothing else, so the context you choose is the whole decision, financially and practically.

Pick one guard. Shrink what it sees. Set its threshold from your own labeled cases. Then decide whether it belongs on this endpoint or the regular one, with the numbers in front of you instead of the marketing.

Sources: OpenAI Decisions API guide, OpenAI API changelog, OpenAI API pricing, Strands Decider 2B announcement.


Medium metadata

  • Title: OpenAI's Decisions API Bills Only What It Reads, and That Changes How You Write Agent Guards
  • Subtitle: No output charges, a higher input rate than the same model elsewhere, and why the context you send a guard is now both its price and its accuracy
  • Tags: OpenAI, AI Agents, LLM, Software Architecture, Machine Learning
  • Canonical: fervorai.dev