Ollama
Ollama adds local decision models for triage and routing
Ollama 0.35.0 adds a /v1/systemone endpoint that runs decision models from Bespoke Labs and Together AI on local hardware.

Ollama can now run decision models on a local machine, small models that answer several typed questions about a piece of text in one request and return choices with probabilities instead of free text.
Ollama released version 0.35.0 on Sept. 28, 2026. Its release notes on GitHub say Ollama “now supports decision models through /v1/systemone, based on TypeSafe’s Jev API.” The release notes name Nimble from Bespoke Labs and Tev1 from Together AI as available models.
One request carries many questions
Decision models take a piece of text and answer a batch of questions about it at once, according to an Ollama blog post dated Sept. 30. The post describes multiple choice, yes/no and score questions. Ollama suggests ticket triage and content moderation as uses.
Ollama lists no additional costs and lower latency when the models run locally. It said Nimble averaged 91 ms per decision on a MacBook Pro with an M5 Max, a figure Ollama measured during real-time gameplay testing. Getting started takes ollama pull nimble and a call through curl or TypeSafe’s Python SDK.
Ollama said it plans MLX acceleration on Apple Silicon and more specialized decision models, including through its cloud service. None of that has shipped.
Nimble is the 9B option
Nimble is fine-tuned from Qwen3.5-9B by Bespoke Labs, according to its Ollama library page, read Oct. 2. It carries an Apache 2.0 license, downloads at 9.3GB to 9.5GB and takes text only, with a 256K-token context and up to 64 questions per request. It needs Ollama 0.35 or later, and the page showed about 20,700 downloads on Oct. 2.
Bespoke Labs reports 75.7% overall accuracy across 13 public datasets and 3,880 decisions. The page lists 81.6% on choice tasks, 80.2% on boolean tasks and 54.6% on score tasks. It gives no eval date.
Tev1 comes in two sizes
Together AI describes Tev1 as an experimental family fine-tuned from Qwen3.5, according to its library page, read Oct. 2. The 4B model is about 4.4GB to 4.5GB and scores 73.3% accuracy in Together AI’s tests. The 0.8B model is about 797MB to 812MB and scores 63.5%. Ollama labels the 0.8B model experimental.
Together AI fine-tuned the family on 37,840 examples from sources that include MultiNLI, BoolQ and Banking77. Its dataset builders and training scripts are MIT licensed. Tev1 shares Nimble’s 64-question limit and 256K context, and the page showed about 11,000 pulls on Oct. 2.
Smaller fixes ship alongside
Settings now opens without waiting for model discovery. Version 0.35.0 also fixes the macOS update menu and icon, which did not show an available update at startup, and MLX model downloads that could stall indefinitely. Requests that include the deprecated typical_p parameter now log a warning instead of failing.
Analysis
We think a classifier that runs on a laptop in under 100 ms and costs nothing per call changes where routing and triage decisions can live. The accuracy figures are the vendors’ own, though. Nimble’s 54.6% on score tasks is the weak spot, and anyone planning to rely on rubric scoring should test it first.
