Data and AI releases

This month658Companies169

658 this month · 169 companies

Dynamic rate limits for serverless retired

Deprecation7 days agoUnverifiedAdded Oct 9, 2026

Serverless no longer applies dynamic rate limits that scale with your recent traffic. Most users will no longer encounter rate limits, though requests can still be limited during periods of high demand. If you get a 429 Too Many Requests response, reduce your request rate and spread out bursts.

Together AI's release notes

More Together AI platform releases

Every Together AI platform release
Betaunverified

Together Link beta

Together Link runs six coding agents on models hosted by Together AI: Claude Code, Codex, OpenCode, and Pi Code in the terminal, plus Claude Desktop (including Cowork) and ChatGPT Desktop. Install it with one command, launch your agent through it, and your normal agent configuration stays untouched. It's now in beta on macOS and Linux.

Updateunverified

What's included:

- Six agents: Launch Claude Code ( tclaude ), Codex ( tcodex ), OpenCode ( topencode ), or Pi Code ( tpi ) in your terminal, or switch Claude Desktop and ChatGPT Desktop to a reversible Together Link profile. OpenCode requires OpenCode 2, and Pi Code requires version 0.80.8 or newer.

Updateunverified

Code sandbox SDK and CLI

The new together-sandbox SDK runs commands and code in isolated runtime environments built from Docker-image snapshots. It ships as a Python SDK, a TypeScript SDK, and a standalone CLI, and is available to organizations on an allowlist ( contact us to request access).

Updateunverified

New dedicated endpoint models

The following models are now available for deployment on dedicated endpoints : hexgrad/Kokoro-82M (text-to-speech). Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice (text-to-speech).

Updateunverified

Higher LoRA rank limit for fine-tuning

You can now train LoRA adapters with a rank of up to 128 for the majority of models, up from 64. The default rank for these models stays at 64, so set lora_r to use a higher one. The model limits response has a new lora_training.default_rank field next to lora_training.max_rank . Run tg fine-tuning model-limits to see both values for a model.

Updateunverified

Batch API files are now retained for 7 days

The input file you upload for a batch job, along with the output and error files the job produces, are now retained for 7 days. After that the files are no longer accessible, so download your results before the window closes. Reusing an uploaded input file across batch jobs also works only within that window. See Batch inference .

Also shipped on Oct 2, 2026

Together AI in October 2026
Previewunverified

You can use Vertical Pod Autoscaler (VPA) with Horizontal Pod Autoscaler (HPA) to automatically optimize...

You can use Vertical Pod Autoscaler (VPA) with Horizontal Pod Autoscaler (HPA) to automatically optimize container CPU requests for workloads that scale replicas based on CPU utilization. This feature is available in Public Preview on clusters running GKE version 1.36.3-gke.1630000 or later. For more information, see Rightsize HPA workloads with VPA .