Ollama 0.40.1
Update3 days agoUnverifiedAdded Oct 10, 2026
- server: proxy cloud usage and balance APIs - llama: fix clef head reads past 2GiB on windows - cmd: remove account step from CLI onboarding - manifest: avoid symlinks on Windows - docs: fix 6 dead links in README community integrations list - docs: fix broken download links in app README - mlx: drop carried metal residency patch now that it is upstream
Topics: LLMs, Streaming, Developer tools
More Ollama releases
Every Ollama releaseOllama 0.40.2
Model upgrades Models downloaded with earlier versions of Ollama are upgraded in the background the first time you run them, for better performance and compatibility when running on llama.cpp. To make downgrading safe, Ollama keeps the original copy as a backup, so upgraded models are temporarily kept on disk.
Ollama 0.35.1
Clef decision models Ollama now supports Clef and Clef Flash, Cloudflare's new open-source decision models, through /v1/systemone. Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.
Ollama 0.35.0
Decision models Ollama now supports decision models through /v1/systemone, based on TypeSafe’s Jev API. Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification.
Ollama 0.40.0
Models run on MLX on Apple Silicon by default In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.
Ollama 0.34.4
- Structured outputs on thinking models now apply in a single pass, making them faster and more reliable. Fixed intermittent "model not found" errors with a large local library Fixed the macOS app becoming unresponsive when checking if ChatGPT or Codex is running. Qwen 3.8 prompt processing is faster on Apple Silicon.
Ollama 0.34.3
GET /api/show now advertises each model's thinking controls and default: Available in the CLI with: ollama show gemma4 thinking levels false, true default true Available in the API with: curl http://localhost:11434/api/show -d ' {"model": "glm-5.3-flash:cloud"} ' { "thinking" : { "values" : [ " low " , " high " , " max " ], "default" : " max " } } Also...
Ollama 0.34.2
- Added first-run setup when running ollama , with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows. Added ollama://apps to open the desktop app’s Apps page directly on macOS and Windows. Fixed excessive memory growth during long generations with MLX speculative decoding. Updated llama.cpp.
Ollama 0.34.1
- MLX safetensors ollama create no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization. - Improved MLX memory handling on Apple Silicon - Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g.
Also shipped on Oct 7, 2026
Ollama in October 2026The microVM sandbox type is now Generally Available (GA) with GKE Sandbox in clusters that run version...
The microVM sandbox type is now Generally Available (GA) with GKE Sandbox in clusters that run version 1.37.0-gke.4713000 and later. MicroVM sandboxes provide hardware virtualization and isolation for untrusted workloads, AI agent runtimes, and multi-tenant environments.
SSH for Cloud Run services and instances is in Preview
SSH for Cloud Run services and instances is in Preview . Use this feature to establish a secure, interactive shell connection to your running instances.
Set credit usage limits and alerts for your workspace
Workspace admins and owners on paid plans can now set usage limits and alerts that keep shared credits under control. Set a credit threshold for the whole workspace, a project, a member, or, on Business and Enterprise plans, a group or an access token, and choose what happens when usage reaches it: an email or in-app alert, or a block that stops building or...
Faster inline text edits
Simple text changes you make with Edit text inline in the preview toolbar now apply in a few seconds. Text that your app builds from data or translates still takes a little longer.
Gemini 3.7 Flash is deprecated for your app's AI features
Google is retiring Gemini 3.7 Flash ( google/gemini-3.7-flash ), so it is now marked deprecated for your app's AI features and stops working on January 28, 2027. Apps that use it keep working until then, and Lovable no longer picks it for new AI features.
The Cost Analysis tab of the Billing page now includes Serverless Inference Live Usage, which lists the...
The Cost Analysis tab of the Billing page now includes Serverless Inference Live Usage, which lists the estimated cost and token counts of your Serverless Inference requests by model and model access key, for both teams and organizations. …