Open-weight decision model

Decisions,
not text.

Bernoulli takes unstructured state and typed questions and returns calibrated probabilities you can branch on. One forward pass per question. No generated tokens, no parsing, no JSON repair, no retries.

~145 msp50 per question
0.025ECE on SST-2
0outbound calls
bernoulli decide1201 ms · 3 questions
state.text
Where is my package? I ordered it two weeks ago and nothing has arrived. This is unacceptable service.
What does the customer want?choice
tracking0.942
other0.041
refund0.015
exchange0.001
Is this urgent?binary
truep = 0.947
How angry is the customer? (1–5)rating
1
2
3
4
5
expected = 4.30
Why

Your app needs a branch, not a paragraph.

The usual way to get a decision from an LLM is to make it write text and then hope the text parses. Bernoulli reads the answer straight out of the model's logits instead.

Prompted LLM generative

  1. 1Prompt: "Reply in JSON…"
  2. 2Generate tokens
  3. 3Parse output, repair malformed JSON
  4. 4Retry on refusal or off-format answer
  5. 5Confidence? Ask the model to guess one

Bernoulli logit scoring

  1. 1State + typed question
  2. 2One forward pass
  3. 3Softmax over label tokens only
  4. 4Debias + calibrate
  5. 5Typed decision with a real probability
Features

Small surface. Strong guarantees.

A·B·C

Three typed questions

choice over 2–26 options, binary yes/no, and rating on an integer scale with full distribution and expected value.

p ≈ 0.9

Calibrated probabilities

Per-question-type temperature scaling fit on held-out data. When Bernoulli says 0.9, it should be right about 90% of the time.

⇄

Position debiasing

Re-scores with the option order reversed (or all cyclic shifts) and averages, cancelling the model's preference for "the first option".

1×

One forward pass

No decoding loop. State goes first and questions last, so many questions over one state share a cached prefix.

⌂

Sovereign by default

Open weights, runs fully on your hardware. No outbound network calls at inference time. Offline mode is a first-class target.

◫

Multimodal state Soon

Text plus up to 8 images per state: screenshots, documents, photos. Long context up to the backbone's native window.

How it works

From state to decision in five steps.

The model never writes an answer. Its next-token distribution, restricted to a handful of single-token labels, is the answer.

01

Build the prompt

…state…
Q: What does the customer want?
A) refund  B) exchange
C) tracking D) other
Answer:

State first, question last, options labeled with letters. The assistant turn is pre-filled so the next token is the answer.

02

One forward pass

logits over ~152k tokens

No decoding loop. Read the scores for every possible next token at the last position.

03

Keep only labels

4 single-token labels

Throw away everything except A–D. Each label is verified at startup to be exactly one token.

04

Debias

orderABCD
reverseABCD
↳ map back · average

Score again with the options reversed, so "tracking" sits in a different slot, then average. Position bias cancels out.

05

Calibrate

÷ T = 1.35
tracking.94
other.04
refund.02
exchange.00

Softmax with a fitted temperature per question type. Out comes a typed decision with an honest probability.

Benchmarks

Same accuracy. Honest confidence.

Against the same backbone prompted to answer in a word and parsed, Bernoulli matches accuracy and gives probabilities you can actually threshold on.

SST-2 · 872 examplesBernoulliGenerative
Accuracy91.74%91.74%
Macro-F10.9170.917
ECE (15 bins) ↓0.0250.083
Brier ↓0.1330.165
NLL ↓0.2342.282
Latency p50 / p95145 / 149 ms136 / 138 ms

Bernoulli row: reverse debiasing + calibrated temperature. Dev backbone Qwen2.5-VL-7B-Instruct (bf16) on a single NVIDIA L4 24GB. Numbers come from scripted eval runs in evals/reports.

3.3×

lower calibration error than parsing the generative answer (ECE 0.025 vs 0.083).

72/72

of the generative baseline's mistakes were stated with 100% confidence. Bernoulli: only 1 of its 72 went above 0.99.

84.8%

zero-shot accuracy on AG News (4-way topic, 7,600 examples) at 155 ms p50.

Know when it's wrong

Same 72 mistakes. Only one shows you where.

Each dot is one SST-2 review, 872 in all. Both methods get the same 800 right. The difference is what the confidence tells you about the other 72.

Prompted LLM generative

Every answer comes back at 1.00. The mistakes look exactly like the correct answers, so there's nothing to filter on.

Bernoulli calibrated

Sorted by confidence (darker = more sure). The mistakes cluster at the low-confidence end, exactly where you'd send them to a human.

correct (shade = confidence) wrong routed to review
20%60%100%
96.1%accuracy on auto-decided
0.83confidence threshold
–mistakes caught for review
–items sent to a human

Data: every prediction from the SST-2 validation run (Qwen2.5-VL-7B-Instruct, reverse debiasing, calibrated). The generative baseline has no usable confidence, so its accuracy stays at about 91.7% however much you hold back.

Interface

One request. Every question answered.

{
  "state": {"text": "Customer email: Where is my package?…"},
  "questions": [
    {"id": "intent", "type": "choice",
     "prompt": "What does the customer want?",
     "options": ["refund", "exchange", "tracking", "other"]},
    {"id": "urgent", "type": "binary",
     "prompt": "Is this urgent?"},
    {"id": "anger", "type": "rating",
     "prompt": "How angry?", "scale": [1, 5]}
  ],
  "options": {"debias": "reverse", "calibrated": true}
}

Typed in, typed out

Pydantic v2 discriminated unions with extra='forbid'. Invalid input is rejected with a 422, never silently guessed.

Your strings come back

Options are scored as letters internally; responses return your original option labels and integer rating keys.

Swappable backbone

A Scorer interface with HF and vLLM backends. Model id and revision are pinned in config, never hard-coded.

Calibrate on your data

Temperatures transfer imperfectly across domains, so domain re-calibration from your own labeled JSONL is planned.

Stop parsing. Start deciding.

Bernoulli is early and moving fast. Star the repo, try the CLI on your own data, and watch it grow.

git clone https://github.com/shyamsfo/bernoulli.git