Together AI develops cloud infrastructure and developer tools for training, fine-tuning, and running open-source AI models. The company focuses on making high-performance generative AI more accessible to startups, researchers, and engineering teams.
Together AI has announced a new three-part resource model for its Dedicated Model Inference service, consisting of endpoints, deployments, and configs. The system incorporates capacity-aware routing to efficiently manage resources and provide more granular control over model deployment configurations.
Why it matters: This update enables more efficient and customizable AI model serving by giving users finer control over dedicated inference resources.
Together AI conducted 452 DeepSWE rollouts comparing Kimi K3 and Claude Fable 5. Claude Fable 5 leads in pass@1 by 1.4 points, while Kimi K3 outperforms in pass@4 and achieves 2.8 times more solves per dollar.
Why it matters: This benchmark offers developers practical insights into the cost-efficiency and coding performance of two leading models.
Together AI has introduced a new production platform for open-weight AI inference. The platform allows users to deploy open models quickly, with control over performance, cost, and quality, and is designed to scale to meet service-level objectives while supporting safe rollout.
Why it matters: This platform provides a managed solution for deploying open-weight models in production, reducing the complexity of building custom inference infrastructure.
Together AI has partnered with Y Combinator to provide a dedicated GPU cluster for YC startups, enabling faster access to compute resources without requiring long-term contracts. This initiative is designed to support AI development within the YC community.
Why it matters: The partnership offers YC startups more flexible and rapid access to GPU resources, which could help accelerate AI innovation among early-stage companies.
Together AI discusses the practical implications of different uptime percentages for inference services, explaining what is required to achieve 99%, 99.9%, and 99.99% availability. The post also describes the types of failures each level must withstand and suggests questions to consider when evaluating inference providers.
Why it matters: Understanding uptime guarantees is important for selecting reliable AI inference providers as these services become integral to production systems.
Thinking Machines Lab has released its first open model, Inkling, a 975-billion-parameter multimodal AI trained to understand video and audio. The model is available on Together AI's platform from day one, serving as the company's first public demonstration after a year and a half of developing AI infrastructure largely out of public view.
Why it matters: Inkling could help position Thinking Machines Lab as a competitor to Anthropic and OpenAI in the open model space.
Together AI announced improvements to its GPU clusters for production AI workloads, including passive health checks, automated node repair, enhanced Slurm reliability, OIDC authentication, and startup scripts. These features are designed to provide greater reliability and control for users running large-scale AI training and inference.
Why it matters: As AI workloads scale, reliable and controllable GPU infrastructure becomes increasingly important for production deployments, and Together AI's updates address key operational challenges.
Together AI has addressed a production bug known as 'Copy Fail,' which was traced back to a 732-byte code change. The company deployed a fix to resolve the issue in their infrastructure.
Why it matters: This highlights how even small code changes can lead to significant issues in AI infrastructure.
Together AI published a guide on designing multi-tenant GPU clusters that pool capacity while maintaining team isolation. The article explains how AI-native companies can achieve this balance and describes Together AI's practical implementation.
Why it matters: This guide provides practical insights for AI teams needing efficient GPU resource sharing without compromising isolation.
Together AI has announced a new technique called distribution-aware speculative decoding (DAS) that can speed up reinforcement learning (RL) rollouts by up to 50% without degrading reward quality. The method addresses the bottleneck of rollout generation in RL post-training by adaptively applying speculative decoding. The announcement was made on the Together AI blog.
Why it matters: This advancement could significantly reduce the time and cost of RL post-training, making it more practical for large-scale AI model development.
Together AI has partnered with Adaption to bring Together Fine-Tuning natively into the Adaptive Data platform. This integration enables teams to optimize datasets, run fine-tuning, evaluate results, and deploy stronger open models.
Why it matters: This partnership streamlines the fine-tuning workflow for open models, making it easier for teams to improve model performance directly from their data platform.
Together AI has achieved ISO 27001:2022 certification, validating its information security management system for enterprise-grade security in production AI workloads. This milestone demonstrates the company's commitment to maintaining high security standards.
Why it matters: The certification assures enterprises that Together AI meets internationally recognized security standards, which may encourage broader adoption of its AI infrastructure.
Together AI published real-world inference benchmarks for coding agents, reporting 31% higher throughput than TensorRT-LLM, 2× better time-to-first-token at saturation, and 76% lower cost than Claude Opus 4.6. The benchmarks focus on scaling inference for agentic coding workloads.
Why it matters: This demonstrates significant performance and cost improvements for deploying coding agents at scale, which could accelerate adoption of AI-assisted development.
Together AI has made NVIDIA Nemotron 3 Super and Nemotron 3 Nano Omni available on its platform. Nemotron 3 Super offers efficient multi-agent reasoning and a 1M-token context window, while Nemotron 3 Nano Omni is a single open model that can process video, images, audio, and text for agentic workloads at scale.
Why it matters: These launches provide developers with production-grade, multimodal AI models optimized for agentic reasoning and scalable deployment.
Together AI has announced an integration with Goose that allows users to deploy any Hugging Face model in a single session using Dedicated Container Inference. This approach removes setup complexity, enabling models to run in a production-grade GPU environment immediately upon release.
Why it matters: This integration streamlines AI model deployment, making it more accessible to developers without requiring infrastructure expertise.
Together AI has launched DeepSeek-V4 Pro, featuring a 512K context length and controllable reasoning modes. The model offers cached-input pricing for long-context workloads such as code agents, document intelligence, and research synthesis.
Why it matters: This release provides developers with a powerful, cost-efficient model for complex reasoning tasks requiring extended context.
Together AI published a blog post detailing how it serves MiniMax-M3 efficiently, enabling 1M-token context and multimodality. The optimizations include KV-block-major sparse attention, paged MSA decode, optimized index scoring, and a Rust-based multimodal gateway.
Why it matters: This demonstrates practical techniques for deploying large multimodal models with long context windows, which is critical for enterprise applications requiring processing of extensive documents and multiple data types.
Together AI published a blog post detailing the inference systems work required to serve DeepSeek-V4, which supports million-token context. The post covers compressed KV layouts, prefix caching, kernel maturity, and endpoint profiles for long-context workloads on NVIDIA HGX B200 hardware.
Why it matters: This highlights the growing importance of inference infrastructure as models scale to million-token contexts, a key challenge for enterprise AI deployment.
Together AI has announced an $800 million Series C funding round aimed at accelerating the shift to open-source AI. The company emphasized that the economics of closed models do not scale and shared plans for future development.
Why it matters: This significant investment highlights growing market confidence in open-source AI as an alternative to proprietary systems.