New dedicated endpoint models
Update8 days agoUnverifiedAdded Oct 9, 2026
The following models are now available for deployment on dedicated endpoints : hexgrad/Kokoro-82M (text-to-speech). Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice (text-to-speech).
More Together AI platform releases
Every Together AI platform releaseTogether Link beta
Together Link runs six coding agents on models hosted by Together AI: Claude Code, Codex, OpenCode, and Pi Code in the terminal, plus Claude Desktop (including Cowork) and ChatGPT Desktop. Install it with one command, launch your agent through it, and your normal agent configuration stays untouched. It's now in beta on macOS and Linux.
What's included:
- Six agents: Launch Claude Code ( tclaude ), Codex ( tcodex ), OpenCode ( topencode ), or Pi Code ( tpi ) in your terminal, or switch Claude Desktop and ChatGPT Desktop to a reversible Together Link profile. OpenCode requires OpenCode 2, and Pi Code requires version 0.80.8 or newer.
Code sandbox SDK and CLI
The new together-sandbox SDK runs commands and code in isolated runtime environments built from Docker-image snapshots. It ships as a Python SDK, a TypeScript SDK, and a standalone CLI, and is available to organizations on an allowlist ( contact us to request access).
New provisioned throughput models
The following models are now available on provisioned throughput : zai-org/GLM-5.3 . zai-org/GLM-5.3-Flash . deepseek-ai/DeepSeek-V4.1-Flash .
Dynamic rate limits for serverless retired
Serverless no longer applies dynamic rate limits that scale with your recent traffic. Most users will no longer encounter rate limits, though requests can still be limited during periods of high demand. If you get a 429 Too Many Requests response, reduce your request rate and spread out bursts.
Higher LoRA rank limit for fine-tuning
You can now train LoRA adapters with a rank of up to 128 for the majority of models, up from 64. The default rank for these models stays at 64, so set lora_r to use a higher one. The model limits response has a new lora_training.default_rank field next to lora_training.max_rank . Run tg fine-tuning model-limits to see both values for a model.
Longer fine-tuning context for Qwen 27B models
Qwen/Qwen3.8-27B , Qwen/Qwen3.6-27B , and Qwen/Qwen3.5-27B now support a 131,072-token context for SFT (up from 32,768) and 65,536 for DPO (up from 16,384), for both LoRA and full fine-tuning.
Batch API files are now retained for 7 days
The input file you upload for a batch job, along with the output and error files the job produces, are now retained for 7 days. After that the files are no longer accessible, so download your results before the window closes. Reusing an uploaded input file across batch jobs also works only within that window. See Batch inference .
Also shipped on Oct 1, 2026
Together AI in October 2026Pricing
Per image, by resolution : resolution Price 768sq $0.041 1k $0.048 2k $0.100 4k $0.607
The Agent Development Kit (ADK), including the gradient-adk Python package and the gradient CLI, is deprecated
The Agent Development Kit (ADK), including the gradient-adk Python package and the gradient CLI, is deprecated. You cannot deploy any new agents with ADK. Agents already deployed with ADK continue to run normally and you are billed for …
DigitalOcean Insights is now available in public preview, providing dashboards, metrics, logs, traces, and...
DigitalOcean Insights is now available in public preview, providing dashboards, metrics, logs, traces, and metric alerts for supported DigitalOcean resources and regions. For more information, see DigitalOcean Insights.
The DigitalOcean API now supports querying Insights metrics and logs and managing alert rules and...
The DigitalOcean API now supports querying Insights metrics and logs and managing alert rules and notification channels in public preview. For available operations and endpoint details, see the DigitalOcean Insights API Reference.
AI Search is generally available
AI Search is now generally available. Usage-based billing begins on November 1, 2026, with included monthly ingestion, storage, semantic query, and full-text query usage. Cloudflare will send a reminder email the week before billing begins. Refer to Limits & pricing for rates and included usage.
Artifacts, Workers - Artifacts is now in open beta
Artifacts , Cloudflare's versioned file system that speaks Git, is now in open beta. Artifacts is built for scale, so you can create a repository per project, user, session, or task. With Artifacts, you can: Deploy repositories to Workers , Connect an Artifacts repository through Workers Builds .