Strands Labs released Strands Decider 2B on 1 October 2026 for decisions inside agent workflows. The release announcement points to open code and a downloadable model checkpoint that developers can inspect and run locally. The useful question is how this particular model selects an answer, and what its published artifacts let a team verify.

A pointer head replaces text generation

According to the project repository, the reference v19 model has 1.9 billion parameters. It starts with a Qwen3.5-2B-Base model, removes the language-model head and uses a small pointer head to score options supplied with a question. A rank-16 LoRA adapter tunes the base. The model compares the hidden state at the answer position with each option's final-token state, so it scores choices in one forward pass instead of generating a reply word by word.

The checkpoint's model card identifies the published files as an adapter and readout head; the base weights are obtained separately. The interface supports yes/no questions, choices among named options and scores on an ordered rubric. It cannot write a prose explanation, and choosing from permitted options does not ensure that the chosen option is correct.

Open artifacts have reproduction limits

The repository publishes training and evaluation scripts and a source inventory. Its data record says generated rows are committed, while public datasets are downloaded rather than redistributed. It also notes that one rebuild from raw exports is not byte-identical. The model card says the base-model revision was inferred rather than pinned by the training hosts. These records make the work inspectable, but the open files alone do not promise an exact recreation of the historical run.

Test thresholds on the intended task

Strands reports 167 correct answers out of 231 public JevBench tasks for its v19 checkpoint and a 115-millisecond median per question on an RTX 3090 under WSL2. Those are team-run measurements on a particular set and hardware, not independent evidence of performance in a production workflow. The model card also warns that changing a question can leave an answer unchanged, long multistep documents are weaker, and its confidence bands were established only on held-out short classification tasks.

Before using the model to gate a tool call, a team should test labelled cases from its own workflow, measure both wrong approvals and wrong denials, and send uncertain or high-cost cases to review. The project's example uses hand-picked questions and a hand-picked threshold. Application code must still own the action decision; a model score is evidence for that decision, not a policy enforcement guarantee.