Signal

Build your own decision model

First reported by Nishtahir ·

The signal ●●○○ Compiled by AI from Nishtahir, Hacker News and The Decoder
Why you might care

The ability to constrain LLM outputs to specific choices reduces the computational cost of generating answers.

What happened

A developer named Nish Tahir has demonstrated how to build a "decision model" using a language model (LLM) that responds with calibrated probabilities for a fixed set of answers. Traditional LLMs generate text token by token, requiring multiple passes. Decision models, however, aim for a single pass by constraining the output to a predefined set of options, like A, B, C, D, or E. This approach uses the LLM to emit only tokens corresponding to these options, selecting the highest probability output. Tahir's example uses the Qwen/Qwen3-1.7B model and shows how to constrain token outputs. He tested this method on a commonsense question dataset, achieving 72.5% accuracy before fine-tuning and 76.2% after. He also highlighted an issue with overconfidence, where models assign high probabilities to incorrect answers, and demonstrated temperature scaling as a method to calibrate these probabilities closer to their actual accuracy.

What it means

The core innovation is transforming a free-form text generation task into a classification problem. By forcing an LLM to choose only from a discrete set of predefined answers, the inference process can be significantly optimized. This is achieved by masking the model's vocabulary to only allow specific tokens that represent the answer choices. The model then outputs probabilities for these limited options, effectively acting as a classifier rather than a pure generator.

A key challenge addressed is model calibration, where the model's stated confidence (output probability) doesn't align with its actual accuracy. Tahir demonstrates that LLMs can be overconfident, especially on ambiguous questions. Techniques like temperature scaling, a method to adjust the "sharpness" of the probability distribution, are shown to improve calibration, making the model's confidence scores more reliable indicators of correctness.

AI-written summary. May contain errors.

Build