Bernoulli takes unstructured state and typed questions and returns calibrated probabilities you can branch on. One forward pass per question. No generated tokens, no parsing, no JSON repair, no retries.
Where is my package? I ordered it two weeks ago and nothing has arrived. This is unacceptable service.
The usual way to get a decision from an LLM is to make it write text and then hope the text parses. Bernoulli reads the answer straight out of the model's logits instead.
choice over 2–26 options, binary yes/no, and rating on an integer scale with full distribution and expected value.
Per-question-type temperature scaling fit on held-out data. When Bernoulli says 0.9, it should be right about 90% of the time.
Re-scores with the option order reversed (or all cyclic shifts) and averages, cancelling the model's preference for "the first option".
No decoding loop. State goes first and questions last, so many questions over one state share a cached prefix.
Open weights, runs fully on your hardware. No outbound network calls at inference time. Offline mode is a first-class target.
Text plus up to 8 images per state: screenshots, documents, photos. Long context up to the backbone's native window.
The model never writes an answer. Its next-token distribution, restricted to a handful of single-token labels, is the answer.
State first, question last, options labeled with letters. The assistant turn is pre-filled so the next token is the answer.
No decoding loop. Read the scores for every possible next token at the last position.
Throw away everything except A–D. Each label is verified at startup to be exactly one token.
Score again with the options reversed, so "tracking" sits in a different slot, then average. Position bias cancels out.
Softmax with a fitted temperature per question type. Out comes a typed decision with an honest probability.
Against the same backbone prompted to answer in a word and parsed, Bernoulli matches accuracy and gives probabilities you can actually threshold on.
| SST-2 · 872 examples | Bernoulli | Generative |
|---|---|---|
| Accuracy | 91.74% | 91.74% |
| Macro-F1 | 0.917 | 0.917 |
| ECE (15 bins) ↓ | 0.025 | 0.083 |
| Brier ↓ | 0.133 | 0.165 |
| NLL ↓ | 0.234 | 2.282 |
| Latency p50 / p95 | 145 / 149 ms | 136 / 138 ms |
Bernoulli row: reverse debiasing + calibrated temperature. Dev backbone Qwen2.5-VL-7B-Instruct (bf16) on a single NVIDIA L4 24GB. Numbers come from scripted eval runs in evals/reports.
lower calibration error than parsing the generative answer (ECE 0.025 vs 0.083).
of the generative baseline's mistakes were stated with 100% confidence. Bernoulli: only 1 of its 72 went above 0.99.
zero-shot accuracy on AG News (4-way topic, 7,600 examples) at 155 ms p50.
Each dot is one SST-2 review, 872 in all. Both methods get the same 800 right. The difference is what the confidence tells you about the other 72.
Every answer comes back at 1.00. The mistakes look exactly like the correct answers, so there's nothing to filter on.
Sorted by confidence (darker = more sure). The mistakes cluster at the low-confidence end, exactly where you'd send them to a human.
Data: every prediction from the SST-2 validation run (Qwen2.5-VL-7B-Instruct, reverse debiasing, calibrated). The generative baseline has no usable confidence, so its accuracy stays at about 91.7% however much you hold back.
{
"state": {"text": "Customer email: Where is my package?…"},
"questions": [
{"id": "intent", "type": "choice",
"prompt": "What does the customer want?",
"options": ["refund", "exchange", "tracking", "other"]},
{"id": "urgent", "type": "binary",
"prompt": "Is this urgent?"},
{"id": "anger", "type": "rating",
"prompt": "How angry?", "scale": [1, 5]}
],
"options": {"debias": "reverse", "calibrated": true}
}
# single or batched questions over one state uv run bernoulli decide \ --state state.txt \ --question questions.json # or a full DecideRequest payload uv run bernoulli decide --request request.json
# FastAPI over the vLLM backend, coming soon curl -s localhost:8000/v1/decide \ -H 'content-type: application/json' \ -d @request.json # also: GET /healthz, GET /v1/models
Pydantic v2 discriminated unions with extra='forbid'. Invalid input is rejected with a 422, never silently guessed.
Options are scored as letters internally; responses return your original option labels and integer rating keys.
A Scorer interface with HF and vLLM backends. Model id and revision are pinned in config, never hard-coded.
Temperatures transfer imperfectly across domains, so domain re-calibration from your own labeled JSONL is planned.
Bernoulli is early and moving fast. Star the repo, try the CLI on your own data, and watch it grow.