What changed in AI — Page 79

Companies & FundingReportedTechCrunch / AI

Applied Computing raises $20M to build AI foundation model for oil and gas plants

Applied Computing has raised a $20M Series A to develop a foundation AI model for the oil, gas, and petrochemical industry. The company aims to provide operators with an AI system designed for entire plant operations.

Why it matters: This funding highlights increasing investment in specialized AI models for heavy industry, which could improve operational efficiency in oil and gas.

ResearchOfficialMIT News / Artificial Intelligence

A better way to turn 2D designs into 3D models for rapid prototyping

MIT researchers have developed an automated framework that helps AI models generate CAD programs from 2D designs more accurately and efficiently. This advancement could make it easier to convert 2D sketches into 3D models for rapid prototyping.

Why it matters: The framework could accelerate product design cycles by reducing the manual effort required to create 3D CAD models from 2D concepts.

ModelsOfficialarXiv Statistical ML

Algorithms Achieve Fixed-Parameter Tractability for Differentially Private Synthetic Data Generation

A new preprint establishes that generating synthetic data under differential privacy is fixed-parameter tractable (FPT) when parameterized by the treewidth of the query family's incidence graph. The authors introduce two algorithms that achieve optimal error rates: one based on linear programming and the FPT of the LP dual's separation problem, and another using a subsampled private multiplicative weights method with FPT Gibbs sampling. Both approaches are unified by a dynamic programming framework over tree decompositions.

Why it matters: This result advances the theoretical understanding of private synthetic data generation, potentially enabling more efficient privacy-preserving data analysis for complex query families.

ResearchOfficialarXiv Statistical ML

Fast Rates for Semi-Supervised Learning via Data-Augmentation Graph Regularization

A new theoretical analysis demonstrates that data augmentation in semi-supervised learning can achieve a fast O(1/n_L) error rate with respect to the number of labeled samples, improving over the standard supervised O(1/√n_L) rate. The bound explicitly connects error to the quality of augmentations, measured by the graph-cut mass of augmentations crossing label boundaries. This work provides a mechanistic explanation for how augmentation quality influences the trade-off between accuracy and label count.

Why it matters: This is the first theoretical result to explain the labeled-sample efficiency of self-supervised learning, offering insights that could help reduce annotation costs in practice.

ResearchOfficialarXiv Statistical ML

Hybrid Model Combines Differentiable PDE Solvers and Neural Networks for Sparse Measurement Reconstruction

A new hybrid modeling pipeline integrates Radial Basis Function reconstruction, a neural network correction, and a differentiable partial differential equation (PDE) solver to reconstruct dense physical fields from sparse measurements. The approach enables training the neural network without access to fully-resolved simulation states, by embedding the differentiable PDE solver directly in the training loop. Evaluated on fluid mechanics benchmarks, the method outperforms existing statistical and machine-learning-based reconstruction techniques.

Why it matters: This method allows for physics-informed reconstruction from sparse data without requiring complete simulation examples, addressing a common limitation in real-world applications.

ResearchOfficialarXiv Statistical ML

Generalised Exponentiated Gradient Algorithm Enhances Fairness in Multi-Class and Binary Classification

Researchers have introduced a Generalised Exponentiated Gradient (GEG) algorithm designed to improve fairness in both binary and multi-class classification tasks. The in-processing method can address multiple fairness constraints simultaneously and was empirically evaluated against six baseline methods on seven multi-class and three binary datasets, using several effectiveness and fairness metrics.

Why it matters: This work advances fairness mitigation techniques by providing a method applicable to multi-class classification, an area that has received less attention despite its growing importance in real-world AI applications.

InfrastructureReportedThe New York Times / AI

India Is Moving Fast to Build A.I. Data Centers. A Coastal City May Pay the Price.

India is rapidly constructing AI data centers in an effort to catch up in the technology sector. However, critics caution that these large-scale projects will consume significant amounts of energy and water, while offering limited long-term employment opportunities. The coastal city where these centers are being built may face substantial environmental impacts.

Why it matters: This underscores the conflict between advancing AI infrastructure and maintaining environmental sustainability in developing regions.

ResearchOfficialarXiv Software Engineering

UniCode: Augmenting Evaluation for Code Reasoning

UniCode is a generative evaluation framework designed to systematically probe the code reasoning abilities of large language models (LLMs). It introduces multi-dimensional augmentation of coding problems, automated test generation, and fine-grained metrics to reveal model limitations. Experiments show that state-of-the-art models experience a 31.2% performance drop on UniCode, mainly due to weaknesses in conceptual modeling and scalability reasoning.

Why it matters: UniCode demonstrates that current coding benchmarks may overstate LLM capabilities by allowing reliance on statistical shortcuts rather than genuine reasoning.

ResearchOfficialarXiv Software Engineering

Inference Economics of Enterprise Coding Agents: Cloud vs. On-Premise LLMs

A longitudinal case study compares API-based Claude Opus with on-premise GLM-5.1/5.2 for enterprise coding agents over two 28-day periods. Prompt caching achieved a 99.3% hit rate, reducing API costs by 88.6% to $0.57 per million tokens, which is lower than the amortized unit cost of the on-premise solution. While on-premise deployment saved 40.1% of total cost of ownership (TCO) under shared GPU allocation, it resulted in a higher defect-repair burden, with a Fix Commit Ratio of 74.9% versus 45.9% for the API-based approach.

Why it matters: This study provides empirical data on the cost and quality trade-offs between cloud and on-premise LLM deployments for enterprise coding agents, offering practical insights for infrastructure decision-making.

ResearchOfficialarXiv Software Engineering

SemaDiff: Identifying Semantic-Changing Commits with Generated Code and Tests

SemaDiff is a novel approach that uses large language model (LLM)-generated code and tests to distinguish between semantic-preserving and behavior-changing commits in software repositories. By generating additional dependent classes and tests for both pre- and post-commit code versions, SemaDiff can detect behavioral differences. In evaluation on a manually annotated dataset of 183 Java commits, SemaDiff achieved 76% accuracy and 100% precision in identifying semantic-changing commits.

Why it matters: This method advances software repository mining by enabling more reliable identification of purely refactoring commits, which is important for tasks such as debugging, fault localization, and constructing bug datasets.

ResearchOfficialarXiv Software Engineering

Design-Specification Tiling Improves In-Context Learning for CAD Code Generation

A new method called Design-Specification Tiling (DST) is proposed for selecting in-context learning exemplars in CAD code generation tasks. DST frames exemplar selection as a submodular maximization problem, providing a (1-1/e)-approximation guarantee, and aims to maximize coverage of design requirements. Experiments across multiple large language models show that DST leads to substantial improvements in CAD code generation quality compared to existing selection strategies.

Why it matters: This work introduces a principled approach to exemplar selection that enhances the effectiveness of LLMs in complex, domain-specific code generation tasks like CAD.

ResearchOfficialarXiv Software Engineering

VisualRepair: Dynamic Tool Calling and Region Focusing for Visual Software Issue Repair

VisualRepair is a new framework for automated program repair that leverages multimodal large language models (MLLMs) to incorporate visual information from bug screenshots. It introduces image type-aware tool calling and dynamic region focusing to improve fault localization and patch generation. On the SWE-bench Multimodal benchmark, VisualRepair resolves 196 test set instances, outperforming the best baseline by 10 instances.

Why it matters: This work demonstrates a meaningful advance in automated program repair by effectively integrating visual information from bug reports, addressing challenges in modern software with graphical interfaces.

ResearchOfficialarXiv Software Engineering

GameEngineBench: New Benchmark Tests Coding Agents on Real C++ Game Engine Tasks

Researchers have introduced GameEngineBench, a benchmark designed to evaluate coding agents on C++ implementation tasks within Unreal Engine 5 projects. The benchmark comprises 110 tasks sourced from nine real-world game repositories, covering a wide range of gameplay and engine-related challenges. The best-performing model achieved only 55.5% pass@1, and 31 tasks were not solved by any tested configuration, underscoring the difficulty of these tasks for current AI coding agents.

Why it matters: This benchmark highlights a significant gap in the ability of current coding agents to handle complex, integrated C++ development in real-time interactive environments, pointing to important limitations in AI-assisted software engineering.

ResearchOfficialarXiv Software Engineering

Monty: Faithful Autoformalization of Natural Language Assertions

Researchers introduce Monty, a framework that leverages large language models (LLMs) to automatically formalize natural-language assertions into executable formal contracts. Monty improves precision by up to 20 points over naive LLM translation by filtering candidate formalizations using conformance and validity scores. The system was evaluated on 541 assertion-generation tasks from 22 Java collection-like classes.

Why it matters: Automating the creation of formal contracts could reduce the manual effort required for software testing and verification, potentially enhancing developer productivity and software reliability.

ResearchOfficialarXiv Software Engineering

VulWeaver: LLM-Based Vulnerability Detection Achieves State-of-the-Art Results

VulWeaver is a novel approach that combines large language models (LLMs) with structured program analysis to improve vulnerability detection in source code. It achieves F1-scores of 0.76 on Java and 0.78 on C/C++ datasets, outperforming existing baselines by 17-25%. In real-world tests, VulWeaver detected 26 true vulnerabilities across 9 Java projects, with 15 confirmed by developers and 5 CVEs assigned.

Why it matters: VulWeaver demonstrates that integrating LLMs with program analysis can significantly advance automated vulnerability detection, with practical impact validated by developer confirmations and CVE assignments.

ResearchOfficialarXiv Software Engineering

PROBE: A New Benchmark for Comprehensive Code Generation Evaluation in LLMs

Researchers introduce PROBE, a benchmark framework for evaluating code generation by large language models (LLMs) across multiple dimensions: functional correctness, proximity to valid solutions, and code quality. PROBE assesses models using five programming languages and three prompting strategies, and includes analysis of common errors. The results show that LLMs struggle with more difficult problems and that smaller models perform worse on low-resource languages, often making fundamental errors.

Why it matters: PROBE enables a more thorough and systematic evaluation of LLM-generated code, revealing persistent reliability challenges that are not captured by existing benchmarks.

ResearchOfficialarXiv Software Engineering

SeedSmith: LLM-Driven Seed Synthesis for Directed Fuzzing

SeedSmith is an agentic LLM pipeline designed to generate initial seed inputs for directed fuzzing by emulating a security analyst's workflow. It iteratively explores codebases, resolves indirect calls, identifies crash preconditions, and synthesizes concrete inputs, serving as a fuzzer-agnostic seed generation front-end. In evaluations, SeedSmith enabled fuzzers to achieve crash-time speedups of 11.51x to 14.66x on Magma and to trigger 16 previously unreachable bugs across 10 projects with diverse input formats.

Why it matters: SeedSmith demonstrates a significant advance in automating and improving the efficiency of vulnerability discovery through LLM-driven seed generation for fuzzing.

ResearchOfficialarXiv Software Engineering

Design-System-Aware AI Tools Significantly Accelerate Front-End Development

A controlled experiment at a large Brazilian enterprise found that design-system-aware AI assistance reduced front-end development time by 46.7% to 69.4% across Angular, iOS, and Android stacks. The study also observed increased task completeness and reduced performance variability compared to manual or design-system-only development. Analysis indicated that AI tools helped reduce workflow friction and improved the consistency of design implementation.

Why it matters: This study provides empirical evidence that integrating AI with design systems can substantially improve the speed, consistency, and quality of industrial front-end development workflows.

Policy & SafetyOfficialarXiv Software Engineering

FairCoder: Probing LLM Bias in High-Stakes Decision Making via Coding Tasks

Researchers introduce FairCoder, a benchmark that frames high-stakes decision-making as coding tasks to systematically probe large language model (LLM) bias in domains such as employment, education, and healthcare. The study also proposes FairScore, a new metric that accounts for both refusal behavior and group-level outcome diversity. Experiments on leading LLMs reveal consistent and previously underexplored bias patterns, including a tendency to prioritize applicants from high-income families in college admissions. The work provides a comprehensive framework for evaluating LLM bias in practical decision-making scenarios.

Why it matters: This research highlights the risks of deploying LLMs in real-world decision-making and offers tools for more thorough bias evaluation.

ResearchOfficialarXiv Software Engineering

Self-Improving AI Coding Agents Through Accumulated Behavioral Rules: A Closed-Loop Framework

A new framework enables LLM-based coding agents to persistently learn from human review feedback by codifying accepted review comments as behavioral rules, without requiring model retraining. Deployed across a 35+ service microservices platform, the rule set expanded from 5 to 18 behavioral rules, eliminating recurrence of previously ruled-against error classes. The system shifts human review focus from low-level correctness to higher-level design validation.

Why it matters: This approach demonstrates a practical method for coding agents to continuously improve and reduce repetitive errors in real-world codebases without retraining.