← Back to Together AI Blog

Together AI Blog briefings

Products & AgentsOfficialTogether AI Blog

Together AI Unveils New Inference, Agents, Voice AI, and Open Models at NVIDIA GTC 2026

Together AI announced new launches in inference, agents, voice AI, and open models at NVIDIA GTC 2026. The company also hosted technical sessions led by its research and engineering leaders.

Why it matters: These launches expand Together AI's platform with new capabilities, highlighting ongoing innovation in the AI infrastructure sector.

InfrastructureOfficialTogether AI Blog

Together AI Adds Autoscaling, Observability, and Self-Healing to GPU Clusters

Together AI has rolled out new features for its GPU Clusters, including autoscaling, role-based access control (RBAC), full-stack observability, and self-healing node repair. These enhancements are designed to deliver production-ready GPU infrastructure that scales efficiently and remains resilient for enterprise workloads.

Why it matters: The update addresses enterprise needs for scalable, reliable, and manageable GPU infrastructure for AI workloads.

ResearchOfficialTogether AI Blog

Together AI Introduces FlashAttention-4 with Pipelining and Hybrid Softmax

Together AI has announced FlashAttention-4, a new kernel design that addresses asymmetric hardware scaling by introducing pipelining for maximum overlap, 2-CTA MMA modes to reduce shared memory traffic, and a hardware-software hybrid approach to softmax exponentials. The technique aims to keep pace with GPU throughput outpacing memory bandwidth.

Why it matters: FlashAttention-4 could significantly improve the efficiency of attention mechanisms in large language models, enabling faster training and inference as hardware continues to evolve.

ModelsOfficialTogether AI Blog

Together AI Announces FlashAttention-4, ThunderAgent, and together.compile at AI Native Conf

At the AI Native Conf, Together AI announced FlashAttention-4, ThunderAgent, and together.compile, highlighting advancements in kernels, reinforcement learning, and inference optimization. The company stated that these research developments are being deployed directly to production on its AI Native Cloud.

Why it matters: These announcements demonstrate ongoing innovation in AI infrastructure, with Together AI moving new research into production environments.

InfrastructureOfficialTogether AI Blog

Together AI unveils CPD architecture for up to 40% faster long-context LLM serving

Together AI introduced Cache-aware Prefill–Decode Disaggregation (CPD), a new inference architecture that separates warm and cold workloads. The approach delivers up to 40% higher throughput and significantly reduces time-to-first-token for long-context LLM serving.

Why it matters: This technique addresses a key bottleneck in serving long prompts, enabling faster responses for applications like document analysis and code generation.

Open SourceOfficialTogether AI Blog

Together AI Releases CoderForge-Preview, an Open Dataset for Training Coding Agents

Together AI has announced CoderForge-Preview, an open dataset intended for training efficient coding agents. The dataset is designed to support open-source AI development in code generation and understanding.

Why it matters: This release offers an open resource that could support research and development of coding AI agents and foster collaboration in the field.

ResearchOfficialTogether AI Blog

Together AI: Speech Models Fail 39% on Street Names, Suggests Solution

Together AI's research finds that leading speech models such as Whisper and Deepgram, despite near-human benchmark scores, fail to correctly transcribe street names 39% of the time. The company outlines a proposed solution to address this significant shortcoming.

Why it matters: This research exposes a major limitation in speech AI that could affect critical real-world uses like navigation and emergency response.

ModelsOfficialTogether AI Blog

Together AI unveils Consistency Diffusion Language Models with up to 14.5x faster inference

Together AI has introduced Consistency Diffusion Language Models (CDLM), a post-training method that enables exact block-wise KV caching and reduces the number of refinement steps required. This approach achieves up to 14.5x latency improvements over standard diffusion language models without sacrificing output quality.

Why it matters: This development makes diffusion language models more practical for real-time applications by significantly reducing inference time while maintaining quality.

Products & AgentsOfficialTogether AI Blog

Together AI Launches Dedicated Container Inference for Custom Models

Together AI has introduced Dedicated Container Inference, a production-grade orchestration service for custom AI models. The service delivers 1.4x to 2.6x faster inference compared to standard approaches.

Why it matters: This enables enterprises to deploy custom models with significantly improved performance, reducing latency and cost for AI inference at scale.

ResearchOfficialTogether AI Blog

Study Reveals Distinct 'Knowledge Priors' in LLM Families

New research from Together AI finds that different large language model (LLM) families exhibit distinct default behaviors when given no specific prompt. According to the study, GPT models tend to generate code and math, Llama models favor narratives, DeepSeek often produces religious content, and Qwen outputs exam questions.

Why it matters: Understanding these inherent biases is crucial for deploying LLMs in applications where neutrality is important.

ModelsOfficialTogether AI Blog

Rime Arcana V3 Turbo and Rime Arcana V3 Now Available on Together AI

Together AI has announced that Rime Arcana V3 Turbo and Rime Arcana V3 are now available on its platform. Users can now access these models through Together AI.

Why it matters: This expands Together AI's model offerings with new versions of the Rime Arcana series.

Companies & FundingOfficialTogether AI Blog

Together AI Hires Alon Gavrielov as VP of Infrastructure Strategy

Together AI has appointed Alon Gavrielov as Vice President of Infrastructure Strategy. The company says this hire deepens its commitment to building reliable, efficient, and scalable infrastructure for AI-native teams.

Why it matters: This hire highlights Together AI's focus on strengthening its infrastructure to support AI-native teams.

Products & AgentsOfficialTogether AI Blog

Together Evaluations now supports comparing top commercial APIs vs. open source models

Together AI has updated its Evaluations platform to support benchmarking models from OpenAI, Anthropic, and Google alongside open-source and fine-tuned models. Users can now compare quality, cost, and performance across providers within a single platform.

Why it matters: This enables data-driven model selection by allowing direct comparison of proprietary and open-source models on the same evaluation platform.

ModelsOfficialTogether AI Blog

Fine-tuned Open-Source LLM Judge Outperforms GPT-5.2 at 15x Lower Cost

Together AI fine-tuned the open-source GPT-OSS 120B model using Direct Preference Optimization on 5,400 preference pairs. The resulting model outperformed GPT-5.2 in human preference alignment for evaluating model outputs, while offering 15x lower cost and 14x faster inference speeds.

Why it matters: This shows that open-source models can surpass proprietary models in specific evaluation tasks with significantly reduced cost and latency.

Open SourceOfficialTogether AI Blog

Together AI Launches DSGym: A Framework for Training Data Science Agents

Together AI has introduced DSGym, a holistic framework for evaluating and training large language model (LLM)-based data science agents. DSGym features over 90 bioinformatics tasks, 92 Kaggle competitions, and synthetic trajectory generation. Together AI reports that their 4B model achieves state-of-the-art performance among open-source models.

Why it matters: DSGym offers a comprehensive benchmark and training environment for data science agents, which could accelerate advancements in automated data analysis.