Dynamic rate limits for serverless retired
Deprecation7 days agoUnverifiedAdded Oct 9, 2026
Serverless no longer applies dynamic rate limits that scale with your recent traffic. Most users will no longer encounter rate limits, though requests can still be limited during periods of high demand. If you get a 429 Too Many Requests response, reduce your request rate and spread out bursts.
More Together AI platform releases
Every Together AI platform releaseNew provisioned throughput models
The following models are now available on provisioned throughput : zai-org/GLM-5.3 . zai-org/GLM-5.3-Flash . deepseek-ai/DeepSeek-V4.1-Flash .
Together Link beta
Together Link runs six coding agents on models hosted by Together AI: Claude Code, Codex, OpenCode, and Pi Code in the terminal, plus Claude Desktop (including Cowork) and ChatGPT Desktop. Install it with one command, launch your agent through it, and your normal agent configuration stays untouched. It's now in beta on macOS and Linux.
What's included:
- Six agents: Launch Claude Code ( tclaude ), Codex ( tcodex ), OpenCode ( topencode ), or Pi Code ( tpi ) in your terminal, or switch Claude Desktop and ChatGPT Desktop to a reversible Together Link profile. OpenCode requires OpenCode 2, and Pi Code requires version 0.80.8 or newer.
Code sandbox SDK and CLI
The new together-sandbox SDK runs commands and code in isolated runtime environments built from Docker-image snapshots. It ships as a Python SDK, a TypeScript SDK, and a standalone CLI, and is available to organizations on an allowlist ( contact us to request access).
New dedicated endpoint models
The following models are now available for deployment on dedicated endpoints : hexgrad/Kokoro-82M (text-to-speech). Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice (text-to-speech).
Higher LoRA rank limit for fine-tuning
You can now train LoRA adapters with a rank of up to 128 for the majority of models, up from 64. The default rank for these models stays at 64, so set lora_r to use a higher one. The model limits response has a new lora_training.default_rank field next to lora_training.max_rank . Run tg fine-tuning model-limits to see both values for a model.
Longer fine-tuning context for Qwen 27B models
Qwen/Qwen3.8-27B , Qwen/Qwen3.6-27B , and Qwen/Qwen3.5-27B now support a 131,072-token context for SFT (up from 32,768) and 65,536 for DPO (up from 16,384), for both LoRA and full fine-tuning.
Batch API files are now retained for 7 days
The input file you upload for a batch job, along with the output and error files the job produces, are now retained for 7 days. After that the files are no longer accessible, so download your results before the window closes. Reusing an uploaded input file across batch jobs also works only within that window. See Batch inference .
Also shipped on Oct 2, 2026
Together AI in October 2026PagerDuty Advance team-level permissions now Generally Available
Building on team-level permissions, Account Admins can now choose which AI agents are enabled for each team. This further empowers teams to roll out AI adoption with confidence. Learn more.
AI Gateway, Web Search API - Introducing Web Search API
Web Search API is now available in beta. Web Search API lets your AI agents and applications search the Internet and ground their responses in live information, instead of guessing URLs or relying on a model's training cutoff. At launch, you can choose between three search providers: Ceramic.ai, Exa, and Linkup .
KV - Workers KV namespace jurisdictions are now generally available
Jurisdictions for Workers KV namespaces are now generally available. When you create a namespace, you can set a jurisdiction to make sure the namespace's data is only durably stored within that region. Jurisdictions can help you comply with data localization regulations such as GDPR or FedRAMP. Supported jurisdictions are eu , us , and fedramp .
Repository security advisory comments API in public preview
You can now read, add, and edit comments on repository security advisories using the REST API, including advisories created from private vulnerability reports. Until now, the discussion on an advisory was only reachable in the web UI, even though it often holds the most useful triage context on a vulnerability report.
Note : Your clusters might not have these versions available
Note : Your clusters might not have these versions available. Rollouts are already in progress when we publish the release notes, and can take multiple days to complete across all Google Cloud zones. Version 1.36.4-gke.1495000 is now the default version for cluster creation in the Rapid channel.
You can use Vertical Pod Autoscaler (VPA) with Horizontal Pod Autoscaler (HPA) to automatically optimize...
You can use Vertical Pod Autoscaler (VPA) with Horizontal Pod Autoscaler (HPA) to automatically optimize container CPU requests for workloads that scale replicas based on CPU utilization. This feature is available in Public Preview on clusters running GKE version 1.36.3-gke.1630000 or later. For more information, see Rightsize HPA workloads with VPA .