Nothing is generated
The answer is one of the options you declared, scored in one pass.
| Needle | LLM call | |
|---|---|---|
| output | 0 tokens | every token billed |
| sampling | none | varies run to run |
| parsing | none | JSON, may retry |
Needle · private preview
A decision model behind one API call. Typed questions in, probabilities out.
Why a forward pass
If a step in your pipeline asks a language model to choose, rate or approve something, it is paying for generation to get a label. Two things change when the label comes from a forward pass instead.
The answer is one of the options you declared, scored in one pass.
| Needle | LLM call | |
|---|---|---|
| output | 0 tokens | every token billed |
| sampling | none | varies run to run |
| parsing | none | JSON, may retry |
Every answer carries its probability. Drag the bar to set where automation stops and a person starts.
3 handled automatically 3 to a person
Where it fits
Customer was charged twice for order 4417 and asks for the duplicate to be refunded.
User: How did last quarter's revenue compare to forecast? Tools: sql, web_search, calculator.
Forum post: Selling 2 concert tickets, DM me. Payment by gift card only, no refunds!!
How it answers
Every question declares its own answer space, so the model can only return an option you defined.
The primitives in the docsSwitching
Needle accepts System One requests. noul works as an alias for claim, and every answer comes back under the name its question used.
Measured in the open
The same requests sent to Needle and to a commercial decision API ("reference" in the charts), checked against known-correct answers after each round of training (the letters E to L).
Needle, stage L +11.8 pts vs reference
Needle, stage K +9.9 pts vs reference
Needle, stage L -24.5 pts vs reference
Needle, stage L -3.9 pts vs reference
$0.20 per million input tokens, the first 10M each month free. Output is always free, because there is none.