Ollama 0.35.1
Update11 days agoUnverifiedAdded Oct 10, 2026
Clef decision models Ollama now supports Clef and Clef Flash, Cloudflare's new open-source decision models, through /v1/systemone. Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.
Topics: LLMs, Developer tools
More Ollama releases
Every Ollama releaseOllama 0.35.0
Decision models Ollama now supports decision models through /v1/systemone, based on TypeSafe’s Jev API. Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification.
Ollama 0.40.0
Models run on MLX on Apple Silicon by default In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.
Ollama 0.34.4
- Structured outputs on thinking models now apply in a single pass, making them faster and more reliable. Fixed intermittent "model not found" errors with a large local library Fixed the macOS app becoming unresponsive when checking if ChatGPT or Codex is running. Qwen 3.8 prompt processing is faster on Apple Silicon.
Ollama 0.34.3
GET /api/show now advertises each model's thinking controls and default: Available in the CLI with: ollama show gemma4 thinking levels false, true default true Available in the API with: curl http://localhost:11434/api/show -d ' {"model": "glm-5.3-flash:cloud"} ' { "thinking" : { "values" : [ " low " , " high " , " max " ], "default" : " max " } } Also...
Ollama 0.40.1
- server: proxy cloud usage and balance APIs - llama: fix clef head reads past 2GiB on windows - cmd: remove account step from CLI onboarding - manifest: avoid symlinks on Windows - docs: fix 6 dead links in README community integrations list - docs: fix broken download links in app README - mlx: drop carried metal residency patch now that it is upstream
Ollama 0.40.2
Model upgrades Models downloaded with earlier versions of Ollama are upgraded in the background the first time you run them, for better performance and compatibility when running on llama.cpp. To make downgrading safe, Ollama keeps the original copy as a backup, so upgraded models are temporarily kept on disk.
Ollama 0.34.2
- Added first-run setup when running ollama , with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows. Added ollama://apps to open the desktop app’s Apps page directly on macOS and Windows. Fixed excessive memory growth during long generations with MLX speculative decoding. Updated llama.cpp.
Ollama 0.34.1
- MLX safetensors ollama create no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization. - Improved MLX memory handling on Apple Silicon - Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g.
Also shipped on Sep 29, 2026
Ollama in September 2026Horizontal scaling for Databricks Apps is now generally available
Horizontal scaling for Databricks Apps is now generally available. You can run a Databricks app across multiple instances behind one app URL for higher availability and concurrency, with zero-downtime deployments and best-effort session affinity. See Horizontal scaling for Databricks apps .
Elastic Cloud Serverless: Another Azure region
Elastic Cloud Serverless is now available in the Microsoft Azure Iowa region. This expansion brings the total number of supported serverless Azure regions to 12.
New Relic Control host-based fleets are now generally available
Linux and Windows host support for Agent Control and Fleet Control reaches General Availability today, matching the GA support for Kubernetes that shipped in September 2025. New Relic Control now gives you one observability control plane across Kubernetes, Linux, and Windows hosts, so your entire fleet is managed the same way no matter where it runs.
Triage for GitLab | Cloud and Self-Hosted | Team, Beta
Triage for GitLab | is now available in beta for GitLab Cloud and self-managed GitLab organizations on the Team plan and higher. The Triage documentation covers access requirements, prioritization, views, and supported actions.
Support for specifying custom target CPU or concurrency utilization using scaling controls is in General...
Support for specifying custom target CPU or concurrency utilization using scaling controls is in General Availability (GA) .
Rich link previews for Agent Runners sites
Every site created with Agent Runners now includes Open Graph tags. When you share a link to your site, it shows up as a proper preview instead of a bare URL. Social media sites and messaging apps read Open Graph tags to build link previews.