In partnership with

Open up almost any AI product and look at what it's asking the model to do.

A lot of it is tiny calls.

❝

Which team gets this support ticket? / Should the agent call a tool right now? / Which model should handle this request? / Is this transaction suspicious?

Every one of those has a handful of possible answers, and the usual way to get one is still heavy: send input to a general model, ask for JSON, parse, validate, then pull out one word.

All of that to get back billing.

At DevDay, OpenAI announced something built for this. It's called the Decisions API, and it's in limited preview right now, with a broader release promised soon.

The idea

You give it the input, a question and a list of allowed answers. It picks one.

What this replaces

Say this ticket comes in:

❝

I upgraded yesterday but my card was charged twice. Can someone refund the second payment?

Most teams would write something like this today:

const response = await openai.responses.create({
  model: "gpt-6-luna",
  input: `Classify as billing, technical,
  account, or sales. Return JSON only.
  ${ticket}`
});

It works.

But you knew all four possible answers before the model ran, and you still made it generate text and hoped the JSON came back fine.

With a decision layer, you just hand over the choices:

const route = await decide({
  input: ticket,
  choices: ["billing", "technical",
            "account", "sales"]
});
// "billing"

Heads up: decide() is my own shorthand.

OpenAI hasn't published the endpoint or request format yet, so be careful with any tutorial showing openai.decisions.create() as working code.

Picking the model (most interesting bit)

OpenAI now has 3 tiers: Luna for focused, high-volume tasks, Sol for complex work at a lower cost, and Astra for the most demanding work.

Most apps I see send everything to the strongest model they can afford.

That gets expensive fast, because Astra costs $10 per million input tokens and Luna costs $0.10. That's a 100x gap, and it's the same on output ($50 vs $0.50).

Put a decision in front instead. It only has to answer one question: is this request simple, complex, or high stakes? Then your app sends it to the right model.

The router's only job is choosing who solves the problem. Your expensive model shows up only when it's needed.

Inside agents

Say an agent can answer, search docs, call an API, or escalate. The typical loop asks the big reasoning model what to do after every single step.

Picking one of four known moves is a much smaller job than planning the whole task. So split them, let the planner plan, and let a decision layer handle the "what next" after each step.

Now your reasoning model spends its time on what needs reasoning. Choosing an action from predefined answers is one of the use cases OpenAI calls out directly.

Once you notice it, it's everywhere

None of these are necessarily easy. You just already know every possible output.

"Can't I just use Structured Outputs?"

You can. GPT-6 Luna already supports them.

Structured Outputs keep a general model's answer in a fixed shape. Decisions starts from a fixed list and asks the model to choose from it.

Whether that ends up cheaper or faster, nobody knows yet because pricing and latency aren't out. Once it's broadly available, benchmark both.

When to skip it

Finding a bug in a repo / figuring out why conversions dropped / researching companies and proposing a strategy - these need exploring and writing, so they stay with an LLM.

The rule I'd use:

❝

Unknown output space → LLM.

Known output space → consider a decision.

One caveat: four possible answers doesn't make a problem easy. Some bounded calls still need serious reasoning.

What I'd build today

You can't build on Decisions yet but you can build the boundary it will plug into:

type Route = "billing" | "technical" | "account" | "sales";

async function decideRoute(
  ticket: string
): Promise<Route> {
  // today: Structured Outputs on Luna
  // later: swap in Decisions
}

The rest of your app just calls decideRoute() and doesn't care how the answer gets made. When access opens up, you change one function.

My take

For the last few years, the LLM call has been the smallest building block in AI software. Every tiny judgment went through it.

Production apps make hundreds of these judgments.

yes / no
A / B / C
tool / answer
continue / stop
cheap model / expensive model

Most of them don't need a model to write anything.

❝

My bet: within a year, the well-built AI apps will have a decision layer sitting in front of their big model, and their API bills will show it.

Use language models when the answer needs to be figured out or written. When you already know every valid answer, hand over the list and let the model pick.

That’s all for today. 

Until tomorrow,
Vaibhav 🤝🏻

If you read till here, you might find this interesting

#AD 1

The AI notetaker that gets the hard words right

Your AI is only as good as what you feed it. Feed it a transcript with your product names, acronyms, and numbers spelled wrong, and you spend the afternoon hand-fixing the output.

Wispr Flow Notetaker uses your dictionary and calendar, so product names, acronyms, numbers, and uncommon names come out spelled right. It pops up when your meeting starts and captures Zoom, Google Meet, Teams, and Slack huddles with one click, no bot joining the call.

Then your meetings show up inside Claude or ChatGPT through the built-in connector. No copy paste. Try Notetaker free on Mac and Windows.

#AD 2

Stop rewriting prompts. Start engineering loops.

Most developers babysit AI one prompt at a time. Top engineers don't, they build loops that plan, execute, and self-correct. The Code built The Ultimate Guide to Loop Engineering, giving you the exact system Silicon Valley engineers use. Sign up and get the guide free, plus a 5-minute daily newsletter.

Reply

Avatar

or to participate