Ollama 0.40.0
Update2 weeks agoUnverifiedAdded Oct 10, 2026
Models run on MLX on Apple Silicon by default In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.
Topics: Vector search
More Ollama releases
Every Ollama releaseOllama 0.34.4
- Structured outputs on thinking models now apply in a single pass, making them faster and more reliable. Fixed intermittent "model not found" errors with a large local library Fixed the macOS app becoming unresponsive when checking if ChatGPT or Codex is running. Qwen 3.8 prompt processing is faster on Apple Silicon.
Ollama 0.35.0
Decision models Ollama now supports decision models through /v1/systemone, based on TypeSafe’s Jev API. Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification.
Ollama 0.34.3
GET /api/show now advertises each model's thinking controls and default: Available in the CLI with: ollama show gemma4 thinking levels false, true default true Available in the API with: curl http://localhost:11434/api/show -d ' {"model": "glm-5.3-flash:cloud"} ' { "thinking" : { "values" : [ " low " , " high " , " max " ], "default" : " max " } } Also...
Ollama 0.35.1
Clef decision models Ollama now supports Clef and Clef Flash, Cloudflare's new open-source decision models, through /v1/systemone. Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.
Ollama 0.34.2
- Added first-run setup when running ollama , with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows. Added ollama://apps to open the desktop app’s Apps page directly on macOS and Windows. Fixed excessive memory growth during long generations with MLX speculative decoding. Updated llama.cpp.
Ollama 0.34.1
- MLX safetensors ollama create no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization. - Improved MLX memory handling on Apple Silicon - Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g.
Ollama 0.40.1
- server: proxy cloud usage and balance APIs - llama: fix clef head reads past 2GiB on windows - cmd: remove account step from CLI onboarding - manifest: avoid symlinks on Windows - docs: fix 6 dead links in README community integrations list - docs: fix broken download links in app README - mlx: drop carried metal residency patch now that it is upstream
Ollama 0.40.2
Model upgrades Models downloaded with earlier versions of Ollama are upgraded in the background the first time you run them, for better performance and compatibility when running on llama.cpp. To make downgrading safe, Ollama keeps the original copy as a backup, so upgraded models are temporarily kept on disk.
Also shipped on Sep 25, 2026
Ollama in September 2026Tags for dashboards, notebooks, Databricks apps, and Genie Agents are now generally available
You can now apply tags to dashboards, notebooks, Databricks apps, and Genie Agents to organize and categorize your workspace assets. This capability is now generally available. Both governed and ungoverned key-value tags are supported. See Apply tags to Unity Catalog securable objects .
Default Enablement of Copilot Features for Copilot Business and Enterprise
We’re introducing a new global default policy for generally available GitHub Copilot features and supported client capabilities in enterprise and organization Copilot settings. For the next 28 days, you can configure this policy, but it won’t affect feature access for your users yet.
2026.09.15
Studio has a redesigned voice catalogue, new voice search controls in Text-to-Speech, and more ways to preview voices in Voice Design. Voice catalogue: Customize layout between tile or list views, group voices by language, and choose which properties to display.
CLI v0.8.1 | CLI
CLI v0.8.1 Continue local work in the cloud: Use cr code handoff --summary to start a new cloud Coding Agent task with your session summary, an optional plan, and an automatically discovered transcript when available. Clearer handoff results: See transfer progress, the branch and commit, transferred filenames, and a direct link to the cloud task.
Free VMs without an account, tracing, OpenCode on Cloud Agents
Your next project can start with an SSH command. This week, we're opening up free VMs with no account required, introducing tracing in Priority Boarding, and bringing OpenCode to Cloud Agents. Let's get into it!
[In preview] Public Preview: Azure HorizonDB supports PostgreSQL 18
Azure HorizonDB is a fully managed, PostgreSQL-compatible, cloud-native database service designed for scalable, high-performance workloads.Azure HorizonDB, now in public preview, includes PostgreSQL 18 support.Learn more.