Data and AI releases

This month801Companies169

801 this month · 169 companies

Ollama 0.40.2

Update2 days agoUnverifiedAdded Oct 10, 2026

Model upgrades Models downloaded with earlier versions of Ollama are upgraded in the background the first time you run them, for better performance and compatibility when running on llama.cpp. To make downgrading safe, Ollama keeps the original copy as a backup, so upgraded models are temporarily kept on disk.

Topics: LLMs, Governance, Developer tools, Performance

Ollama's release notes

More Ollama releases

Every Ollama release
Updateunverified

Ollama 0.40.1

- server: proxy cloud usage and balance APIs - llama: fix clef head reads past 2GiB on windows - cmd: remove account step from CLI onboarding - manifest: avoid symlinks on Windows - docs: fix 6 dead links in README community integrations list - docs: fix broken download links in app README - mlx: drop carried metal residency patch now that it is upstream

Updateunverified

Ollama 0.35.1

Clef decision models Ollama now supports Clef and Clef Flash, Cloudflare's new open-source decision models, through /v1/systemone. Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.

Updateunverified

Ollama 0.35.0

Decision models Ollama now supports decision models through /v1/systemone, based on TypeSafe’s Jev API. Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification.

Updateunverified

Ollama 0.40.0

Models run on MLX on Apple Silicon by default In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.

Updateunverified

Ollama 0.34.4

- Structured outputs on thinking models now apply in a single pass, making them faster and more reliable. Fixed intermittent "model not found" errors with a large local library Fixed the macOS app becoming unresponsive when checking if ChatGPT or Codex is running. Qwen 3.8 prompt processing is faster on Apple Silicon.

Updateunverified

Ollama 0.34.3

GET /api/show now advertises each model's thinking controls and default: Available in the CLI with: ollama show gemma4 thinking levels false, true default true Available in the API with: curl http://localhost:11434/api/show -d ' {"model": "glm-5.3-flash:cloud"} ' { "thinking" : { "values" : [ " low " , " high " , " max " ], "default" : " max " } } Also...

Updateunverified

Ollama 0.34.2

- Added first-run setup when running ollama , with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows. Added ollama://apps to open the desktop app’s Apps page directly on macOS and Windows. Fixed excessive memory growth during long generations with MLX speculative decoding. Updated llama.cpp.

Updateunverified

Ollama 0.34.1

- MLX safetensors ollama create no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization. - Improved MLX memory handling on Apple Silicon - Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g.

Also shipped on Oct 8, 2026

Ollama in October 2026
GAunverified

Performance Risks Inbox is now generally available

Performance Risks Inbox is GA Performance Risks Inbox is now generally available (GA). It finds the performance anti-patterns that quietly slow your application down, such as N+1 queries and inefficient database calls, and groups them in one place so your team can fix them before they become incidents.

Betaunverified

CLI v0.9.0 | CLI, Beta

CLI v0.9.0 Manage cloud tasks with cr code , including interactive messages, side questions, plans, and delivery. Complete shell commands and read richer review JSON . Verify signed Linux releases and recover more reliably from login and update failures. Run cr update . See cloud tasks and upgrade details .