Llama Stack 1.5.0
UpdateyesterdayUnverifiedAdded Oct 10, 2026
Minor release off the new release-1.5.x branch. 108 commits since v1.4.0. Breaking changes Serve conversations and prompts with responses, or when explicitly listed An explicit apis: list is now authoritative, even when empty Raise the FastAPI floor to 0.137.2 and drop legacy route collection Allow an empty choices list in the OpenAI legacy completions...
Topics: LLMs, MCP, Vector search, Postgres, Developer tools
More Llama Stack releases
Every Llama Stack releaseLlama Stack 1.0.4
Patch release on the release-1.0.x maintenance line. Fixes Require a compatible sentence-transformers on 1.0.x Build and release tooling Wait for published wheels before building 1.0 release images Update the 1.0 release lockfile to ogx-client 1.0.3
Llama Stack 1.4.0
- chore: bump fallback_version to 1.3.1.dev0 fix(starter): default trust_remote_code to false for sentence-transformers feat: Add Meta AI to ogx go auto-detection feat: add local_api_key auth provider with attributes and startup validation feat: default ogx go to localhost instead of all interfaces fix(file-processor): Preserve content in async Docling...
Llama Stack 1.0.3
This maintenance release for the 1.0 line addresses JOSS review findings and includes previously backported security, storage, and provider fixes. Enforce configured route restrictions when creating batches, preserving the submitting caller's identity and provider credentials. #6522 Update UI production dependencies to resolve the reported audit findings.
Llama Stack 1.3.1
- chore: update ogx-client to ^1.3.0 in UI lockfile fix(vector-io): persist vector store metadata to kvstore in Milvus, Chroma, and Weaviate (backport #6371 ) by @mergify [bot] in #6421 fix(prompts): raise typed not-found errors to preserve HTTP keep-alive (backport #6440 ) by @mergify [bot] in #6442 fix(vertexai): map "default" service_tier to None instead...
Llama Stack 1.2.5
- chore: update ogx-client to ^1.2.4 in UI lockfile fix(prompts): raise typed not-found errors to preserve HTTP keep-alive (backport #6440 ) by @mergify [bot] in #6441 fix(vertexai): map "default" service_tier to None instead of "standard" (backport #6424 ) by @mergify [bot] in #6445 fix(vertexai): persist Gemini thought_signature across tool-call round...
Llama Stack 1.2.4
Fixes Persist vector store metadata to the configured KV store for Milvus, Chroma, and Weaviate so registered vector stores remain queryable after server restarts and across worker instances. ( #6371 , backport #6420 ) Maintenance Update the UI lockfile to use ogx-client 1.2.3. ( #6407 )
Llama Stack 1.2.3
- fix: make AsyncOgxClient functional by inheriting from sync ApiClient… v1.2.2...v1.2.3
Also shipped on Oct 9, 2026
Meta in October 2026Google Workspace connector is available in Beta
The managed Google Workspace connector is now available in Beta in Databricks Lakeflow Connect. Use the connector to ingest audit activity from Google Workspace applications and services into Databricks. The connector uses OAuth user authorization and supports incremental ingestion.
Agentic Marketplace Discovery (Public preview)
Agentic Marketplace Discovery is now available in public preview. On the public Snowflake Marketplace Discover page in Snowsight, you can switch between Agentic discovery and Marketplace search . Describe what you need to CoCo to find and compare matching listings, or search for providers and listings directly.
All AI Gateway models with prepaid credits, backend features with your coding agent, and new ways to send feedback release - Oct 09, 2026
All AI Gateway models now available with prepaid credits. You don't need to request access to frontier models in the Neon AI Gateway anymore. If you're on a pai...
Use request tags in Unity Gateway service policies (Beta)
Custom service policies in Unity Gateway can now check caller-supplied request tags, such as requiring a tag before a model request or MCP tool call proceeds. See Request tags .
Use OpenJev as the evaluator for LLM-as-a-judge service policies (Beta)
You can now select OpenJev (Qwen3.5 4B) as the evaluator model service of a custom LLM-as-a-judge service policy. OpenJev classifies content without generating output tokens, so it typically responds faster than a chat evaluator. See Use OpenJev as the evaluator .
Code execution in Genie One (Public Preview)
Genie One can now run code in a secure, isolated sandbox to take on work that goes beyond SQL, such as more advanced analyses. Code runs on your behalf using your existing permissions, so Unity Catalog governance and data access controls continue to apply. See Code execution in Genie One .