← Back to RunPod Blog

RunPod Blog briefings

Products & AgentsOfficialRunPod Blog

RunPod Launches One-Click Invoke AI Deployment for Stable Diffusion

RunPod now offers a one-click template to deploy Invoke AI's Stable Diffusion tools, including the infinite canvas feature. The setup requires minimal configuration, making it easier for users to access advanced image generation capabilities.

Why it matters: This simplifies access to advanced AI image generation tools by reducing deployment complexity.

ModelsOfficialRunPod Blog

Training StyleGAN3 with Vision-Aided GAN on RunPod

A new blog post on RunPod discusses how to train StyleGAN3, a generative adversarial network known for high-resolution image generation without aliasing artifacts, using Vision-Aided GAN techniques. The post details the process and benefits of running such training on RunPod's cloud infrastructure.

Why it matters: This highlights practical approaches for developers to train advanced GAN models using cloud resources.

InfrastructureOfficialRunPod Blog

Agentic AI Workflows Explained: Patterns, Infrastructure, and GPU Requirements

RunPod's blog explains that agentic workflows differ from single model calls by planning, looping, and bursting, which affects the underlying infrastructure. The article discusses workflow patterns, infrastructure needs, and GPU requirements for agentic AI systems.

Why it matters: Understanding the infrastructure demands of agentic AI workflows is crucial for developers and enterprises deploying autonomous agents.

ModelsOfficialRunPod Blog

RunPod Blog: Optimizing Mistral-7B Deployment with Quantized GGUF and vLLM

RunPod published a guide on optimizing Mistral-7B deployment using quantized GGUF models and vLLM workers. The article discusses comparing GPU performance across pods and serverless endpoints.

Why it matters: This provides practical optimization techniques for deploying Mistral-7B efficiently on RunPod's infrastructure.

Products & AgentsOfficialRunPod Blog

Runpod and RandomSeed Bring Stable Diffusion API Access

Runpod has partnered with RandomSeed to offer easy-to-use API access for Stable Diffusion via AUTOMATIC1111. This collaboration is designed to make generative art more accessible to developers.

Why it matters: The partnership lowers the barrier for developers to integrate generative art into their applications by simplifying API access to Stable Diffusion.

Companies & FundingOfficialRunPod Blog

ScribbleVet Uses RunPod Infrastructure for AI-Powered Veterinary Care

ScribbleVet leverages RunPod's infrastructure to provide real-time insights and automated diagnostics in veterinary care. A recent case study highlights how these AI-driven tools contribute to improved outcomes for veterinary professionals and their patients.

Why it matters: This demonstrates how specialized AI infrastructure can enable transformative applications in niche healthcare fields like veterinary medicine.

ModelsOfficialRunPod Blog

Mistral AI Releases Mistral Large 3 and Devstral 2 as Open Models

In early December 2025, Mistral AI released Mistral Large 3 and Devstral 2, both under the Apache 2.0 license. Mistral Large 3 is aimed at high-performance AI applications. The models are available as open-source.

Why it matters: Mistral AI's release of two open models under a permissive license strengthens the open-source AI ecosystem and provides developers with powerful, freely available tools.

Open SourceOfficialRunPod Blog

Introduction to vLLM and PagedAttention

vLLM achieves higher throughput than Hugging Face Transformers by using PagedAttention to eliminate memory waste and boost inference. This technique optimizes memory management for large language models, resulting in more efficient deployment.

Why it matters: PagedAttention addresses a key bottleneck in LLM inference, enabling faster and more efficient deployment of large models.

InfrastructureOfficialRunPod Blog

LLM Agents in Production: What Nobody Tells You About GPU Deployment

RunPod's blog discusses the shift from stateless inference to stateful architectures to resolve infrastructure bottlenecks such as memory management, concurrency limits, and runaway jobs in production AI agents. The article highlights common challenges encountered when deploying LLM agents on GPUs.

Why it matters: This provides practical guidance for developers deploying LLM agents at scale, addressing real-world infrastructure issues.

ResearchOfficialRunPod Blog

RunPod Report: Real AI Usage Data Contradicts Popular Narratives

RunPod's State of AI report, drawing on production data from over 500,000 developers, shows that actual AI workloads differ from widely held beliefs. The report highlights which models and tools are being used in real-world production environments.

Why it matters: This data-driven report provides a clearer picture of AI adoption, enabling developers and businesses to base decisions on real usage rather than assumptions.

Products & AgentsOfficialRunPod Blog

RunPod Launches Faster-Whisper Endpoint: 2-4x Faster and Significantly Cheaper Than Original Whisper

RunPod has introduced a new Faster-Whisper serverless endpoint that delivers 2-4x faster transcription speeds compared to the original Whisper API, at a significantly lower cost. The service is aimed at improving efficiency and affordability for speech transcription tasks.

Why it matters: This development makes high-speed, cost-effective speech transcription more accessible for developers and enterprises relying on audio processing.

Open SourceOfficialRunPod Blog

Upscaling Videos Using VSGAN and TensorRT

RunPod published a step-by-step guide for high-speed video upscaling using VSGAN and TensorRT. The guide details model conversion, engine building, and efficient upscaling on RunPod infrastructure.

Why it matters: This guide helps developers leverage TensorRT acceleration for faster video upscaling, improving efficiency in AI-powered video processing workflows.

ModelsOfficialRunPod Blog

Stable Diffusion 3.5 Delivers Major Quality Leap with Photorealism and Easier Prompts

Stable Diffusion 3.5 has been released, offering a significant improvement in image quality, including photorealistic outputs from minimal prompts. The update addresses previous flaws and enhances ease of use.

Why it matters: This release marks a notable advancement in AI image generation, making high-quality photorealism more accessible with simpler prompts.

Open SourceOfficialRunPod Blog

RunPod Publishes Guide to Deploy Llama 3.1 405B with Ollama

RunPod has released a step-by-step guide for deploying Meta's open-source Llama 3.1 405B model using Ollama on its platform. The guide aims to simplify the deployment process for users interested in running large language models.

Why it matters: This guide makes it easier for users to deploy one of the largest open-source language models, expanding access to advanced AI tools.

Products & AgentsOfficialRunPod Blog

RunPod Launches Overdrive to Optimize AI Inference

RunPod has introduced Overdrive, a new optimization tool designed to improve the efficiency of AI inference workloads. The tool aims to help users get more performance out of their existing model deployments.

Why it matters: This tool could reduce inference costs and latency for developers running AI models on RunPod's infrastructure.

Products & AgentsOfficialRunPod Blog

RunPod Flash Now Generally Available for Serverless GPU/CPU Workloads

RunPod has announced the general availability of Flash, a production-ready tool for running serverless GPU and CPU workloads in pure Python without Docker. The tool is designed to simplify deployment and scaling of AI workloads.

Why it matters: This release lowers the barrier for developers to deploy serverless AI workloads by eliminating the need for Docker, potentially accelerating AI application development.

Products & AgentsOfficialRunPod Blog

RunPod Serverless Updates: Faster Cold Starts, Batch Inference, No-Docker Deploys

RunPod has introduced new serverless features, including faster cold starts, support for batch inference, and the option to deploy without Docker. These updates are designed to enhance performance and reduce costs for users running production endpoints.

Why it matters: These enhancements make serverless AI inference more efficient and accessible for developers deploying models at scale.