Skip to main content

Needle · private preview

Stop asking an LLM to pick an option

A decision model behind one API call. Typed questions in, probabilities out.

Join the list. We send API keys as preview capacity opens, plus new research. Privacy

Read the docs
LLM call
{"action": "refund"}
about 2,400 ms
Needle
refund 0.97
deny 0.02
escalate 0.01
41 ms
“Should the duplicate charge be refunded?” Needle: measured median. LLM: a typical small hosted model with structured output, an illustration.
median latency, measured
41 ms
output tokens, ever
0
accuracy on requests neither model had seen, vs 85.8%
97.0%
answers changed when the options are shuffled, vs 8.3%
0%

Why a forward pass

What changes when nothing is generated

If a step in your pipeline asks a language model to choose, rate or approve something, it is paying for generation to get a label. Two things change when the label comes from a forward pass instead.

Nothing is generated

The answer is one of the options you declared, scored in one pass.

NeedleLLM call
output0 tokensevery token billed
samplingnonevaries run to run
parsingnoneJSON, may retry

You decide how sure is sure enough

Every answer carries its probability. Drag the bar to set where automation stops and a person starts.

Threshold at 0.80
  1. Charged twice for order 4417 refund 0.97 → auto → human
  2. Site down for our whole team escalate 0.94 → auto → human
  3. Wrong size, want an exchange exchange 0.88 → auto → human
  4. Promo code did not apply refund 0.76 → auto → human
  5. Cancel and refund, I was told I could refund 0.54 → auto → human
  6. Is the blue one back in stock? answer 0.66 → auto → human

3 handled automatically 3 to a person

See the benchmark

Where it fits

Three requests, every answer

Support triage

Customer was charged twice for order 4417 and asks for the duplicate to be refunded.

action · choice refund 0.97
urgent · claim 0.22
severity · score 0.61 of 2
Agent tool routing

User: How did last quarter's revenue compare to forecast? Tools: sql, web_search, calculator.

tool · choice sql 0.91
clarify · claim 0.12
risk · score 0.07 of 2
Content review

Forum post: Selling 2 concert tickets, DM me. Payment by gift card only, no refunds!!

action · choice flag 0.63
scam · claim 0.86
toxicity · score 0.05 of 2

How it answers

Three question types, one pass

Every question declares its own answer space, so the model can only return an option you defined.

The primitives in the docs
Three primitives
Choice "choice"
refund 0.97
deny 0.03
→ refund confidence 0.94
Score "score"
lowmediumhigh
→ 0.61 confidence 0.22
Claim "claim"
0 false true 1
→ 0.22 P(true)
What each question type returns for one request: the support-triage example. Every answer carries its probabilities; nothing is generated.

Switching

Already on System One? Change one line

Needle accepts System One requests. noul works as an alias for claim, and every answer comes back under the name its question used.

- base_url = "https://<your System One endpoint>" + base_url = "https://api.gaugenumerics.com/v1"

SDK compatibility

Measured in the open

Every result published, including the losses

The same requests sent to Needle and to a commercial decision API ("reference" in the charts), checked against known-correct answers after each round of training (the letters E to L).

Progress by training stage
Blind test 1180 fields
50 60 70 80 90 100 E G I J K L reference 85.8% 89.7% 98.2% 98.6% 97.5% 97.0% 97.6% 97.6%

Needle, stage L +11.8 pts vs reference

New wording and layouts 432 fields
50 60 70 80 90 100 E G I J K L reference 88.2% 97.9% 96.3% 98.1% 98.1%

Needle, stage K +9.9 pts vs reference

Unseen task families 675 fields
50 60 70 80 90 100 E G I J K L reference 96.1% 63.9% 63.6% 80.4% 71.6% 71.6%

Needle, stage L -24.5 pts vs reference

Reordered choices 1008 fields
50 60 70 80 90 100 E G I J K L reference 93.9% 97.0% 97.8% 82.9% 90.0% 90.0%

Needle, stage L -3.9 pts vs reference

Field accuracy on the same held-out panels as each stage finished, against the reference API on identical requests. Solid points come from a published evidence file, hollow ones from the lab note.

Join the preview

$0.20 per million input tokens, the first 10M each month free. Output is always free, because there is none.