← Back to arXiv Computers and Society

arXiv Computers and Society briefings

ResearchOfficialarXiv Computers and Society

Persona Migration and Expectation Recalibration in Generative AI Adoption: A Longitudinal Study at a State Department of Transportation

A longitudinal study of Microsoft 365 Copilot adoption at a state Department of Transportation found that perceived usefulness of the tool declined significantly after an eight-week pilot, while perceived ease of use, behavioral intention, and trust showed only minor, nonsignificant changes. The research identified three baseline user personas—Skeptics, Cautiously Positive users, and Champions—with substantial individual movement between groups, including 68% of Champions shifting to less enthusiastic personas. The study also observed that while some concerns (accuracy, privacy) decreased, worries about job impact and required skills increased.

Why it matters: This study provides empirical evidence that employee enthusiasm for generative AI tools can wane after real-world use, highlighting the need for ongoing expectation management and tailored support in enterprise AI deployments.

ResearchOfficialarXiv Computers and Society

Auditing Asset-Specific Preferences in Financial LLMs: Bitcoin Bias Identified and Manipulated

A new preprint audits nine leading large language models (LLMs) for asset-specific preferences and finds that Bitcoin's ranking among money-like instruments is highly dependent on the scenario presented, with models ranking it higher in crisis and autonomous-agent contexts. The study identifies a dominant internal feature in Gemma 3 that selectively represents Bitcoin; manipulating this feature can causally shift the model's portfolio allocation toward or away from Bitcoin by up to 5.2 percentage points. This demonstrates that internal model representations can be both audited and causally linked to real financial decisions.

Why it matters: This work provides a concrete method for auditing and influencing asset-specific preferences in financial LLMs, laying groundwork for transparency and accountability as such models are deployed in real-world financial decision-making.

ResearchOfficialarXiv Computers and Society

Study Finds LLM Outputs Lack Diversity Compared to Humans, Proposes Simple Interventions

A new preprint finds that large language models (LLMs) generate responses that are more concentrated and mainstream than the diverse, long-tail outputs produced by humans. The researchers tested interventions such as increasing temperature sampling, prompting for diverse perspectives, and aggregating outputs from multiple models, finding that these methods can improve diversity, but single-model outputs still fall short of human-level diversity.

Why it matters: The study raises concerns that LLMs could reduce cultural diversity in generated content, which has implications for AI policy and the preservation of democratic values.

Policy & SafetyOfficialarXiv Computers and Society

NOHARM benchmark reveals severe harm potential in LLM medical advice; human-AI teaming shows promise

A new benchmark, NOHARM, evaluates 20 large language models (LLMs) and 4 clinical AI tools on 1,100 medical consultation cases, finding that direct use of AI-generated recommendations could result in severe harm in up to 24.6% of cases, with omission errors accounting for over 80% of severe errors. In a randomized study of 101 physicians, AI assistance improved performance, but physicians often omitted valuable AI recommendations, indicating complementary strengths in human-AI teaming.

Why it matters: This study provides the first systematic measurement of clinical safety in LLM-generated medical advice, revealing that widely used AI tools can produce potentially harmful recommendations and highlighting the need for explicit safety evaluation.

ResearchOfficialarXiv Computers and Society

EG-VAR: Lean 4-Based Architecture Eliminates LLM Hallucination in Empirical Reasoning

Researchers introduce EG-VAR, a Lean 4-based tool-calling architecture that ensures every verified output is grounded in attested tool calls and kernel-checked inference, thereby eliminating unsupported claims. On a subset of TableBench numerical reasoning tasks (n=120), EG-VAR achieves perfect accuracy (120/120) compared to a 95% baseline, and maintains 100% source-faithfulness on stress tests where baselines drop to 80-90%. The system also provides explicit audit trails for abstentions and formalization errors.

Why it matters: EG-VAR offers a practical and auditable approach to eliminating LLM hallucination in empirical inference, potentially transforming trustworthiness in high-stakes AI applications.

ResearchOfficialarXiv Computers and Society

LLMs Outperform Traditional ML in Open-Ended Survey Analysis but Face Consistency and Explainability Issues

A new preprint compares large language models (LLMs) such as GPT, Twitter-roBERTa, and LLaMA to traditional machine learning methods for analyzing open-ended survey responses. The study finds that LLMs achieve higher classification accuracy, especially in sentiment and thematic analysis, but exhibit significant variation in consistency and the explicitness of their reasoning. These results highlight important trade-offs between predictive performance and interpretability in large-scale qualitative research.

Why it matters: The study offers practical insights for researchers seeking to balance automation with interpretive rigor when applying LLMs to qualitative data analysis.

ResearchOfficialarXiv Computers and Society

AgentSociety 2: An Integrated Research Environment for Executable Social Science

AgentSociety 2 is a new research environment that integrates large language model (LLM) agents as both AI social scientists and simulated participants, automating the end-to-end workflow of social science research. The system enables hypothesis generation, experiment design, simulation execution, result interpretation, and manuscript drafting across micro, meso, and macro social scenarios. It demonstrates the ability to reproduce qualitative patterns from prior studies and supports large-scale, auditable simulations.

Why it matters: This work advances computational social science by providing a scalable, reproducible, and auditable platform for automating complex social experiments with AI agents.

ResearchOfficialarXiv Computers and Society

Policy-as-Prompt Moderation with LLMs: Risks and Governance Considerations

A new preprint examines the use of large language models (LLMs) for content moderation through 'policy-as-prompt' methods, where moderation policies are given to LLMs as natural-language prompts. The authors argue that this approach introduces specific risks and harms, and that simply writing prompts is not sufficient for effective or meaningful community governance. They propose several considerations for improving prompt governance but conclude that prompt-writing alone cannot ensure robust moderation outcomes.

Why it matters: This research highlights important limitations and governance challenges for AI-driven, prompt-based content moderation systems as their use expands in online communities.

ResearchOfficialarXiv Computers and Society

Study Finds Gap Between Institutional and Course-Level GenAI Policies in Computing Education

A new preprint compares institutional and course-level generative AI policies at U.S. research-intensive universities. The study finds that while institutions tend to be more supportive of GenAI use, course-level guidance in computing education remains cautious. The authors propose an instructor-centered framework to guide future GenAI adoption in courses.

Why it matters: This research highlights a disconnect between university-wide AI policies and classroom practices, offering a framework to help computing educators navigate GenAI adoption.

Policy & SafetyOfficialarXiv Computers and Society

AAAI-26 Desk-Rejects 141 Papers Amid Surge in Undisclosed Dual Submissions

AAAI-26 organizers report a significant increase in dual submissions—papers submitted to multiple venues without disclosure—during the conference's review process. By combining similarity assessment, LLM-based overlap tools, and manual review, they desk-rejected 141 main-track submissions. The organizers warn that generative AI may be enabling more sophisticated forms of dual submission and propose several policy and technical recommendations to address the issue.

Why it matters: This development exposes a growing integrity challenge in AI research, with generative AI potentially exacerbating threats to the peer-review process and the reliability of the scientific record.

ResearchOfficialarXiv Computers and Society

Transfer Learning Across Policy Regimes in Adaptive Multi-Agent Systems

A new preprint examines how transfer learning can be applied to adaptive multi-agent systems facing policy regime changes. The authors compare blank-slate learners with transfer learners that reuse structural knowledge from previous regimes, using an emissions-regulation simulation. Results show that transfer learning improves performance when the policy-outcome relationship remains stable, but can lead to negative transfer when a regime change introduces a threshold break. The paper provides a methodological framework for determining when regulatory experience should be reused or discarded.

Why it matters: This work offers a formal approach to understanding the risks and benefits of transfer learning in policy modeling for adaptive socio-technical systems, informing the design of AI-driven regulatory tools.

Policy & SafetyOfficialarXiv Computers and Society

The Benchmark Ceiling: Human Judgment, Evaluation Scarcity, and the Political Economy of AI Capability Measurement

A new preprint argues that as AI models approach top performance on existing benchmarks, the remaining discriminating items are those requiring elite expert judgment, which is structurally scarce. The authors present a formal model showing how benchmark signal depreciates as models improve, and document a scarcity premium for high-judgment evaluation labor. The paper also discusses the governance implications of these findings for AI capability measurement.

Why it matters: This work highlights a fundamental bottleneck in evaluating advanced AI systems, raising concerns about the reliability of current benchmarks and the challenges for AI governance.

ResearchOfficialarXiv Computers and Society

Model Collapse: Recursive Training Degrades AI but Opens Creative Possibilities

A new preprint explores 'model collapse,' the phenomenon where recursively training AI models on AI-generated data leads to degraded performance, such as repetition and noise. The author argues that while this is typically seen as a technical failure, it also has creative and aesthetic dimensions, drawing parallels to early video feedback art. The paper examines how model collapse challenges certain technological ideals and underscores AI's ongoing reliance on human-generated data.

Why it matters: This research reframes model collapse as not only a technical issue but also a source of artistic and philosophical insight, broadening our understanding of AI's creative potential and limitations.

ResearchOfficialarXiv Computers and Society

Large-Scale Empirical Study of Open-Source Image Generation Model Ecosystem

A new empirical study analyzes 6 million images from the open-source image generation ecosystem, examining how creators use 22,400 base models and 154,000 LoRA models. The research identifies the ecosystem's unique strengths and challenges, offering insights that could inform its sustainability and future innovation. The dataset compiled for the study is publicly available for further research and practical use.

Why it matters: This is the first large-scale empirical analysis of creative workflows in the open-source image generation ecosystem, providing foundational insights for researchers and practitioners.

ResearchOfficialarXiv Computers and Society

Return of the Solo Author: The Changing Division of Labor in Science in the Age of Generative AI

A large-scale study analyzing over 300 million scientific works across 26 fields finds that the decades-long decline in solo-authored papers halted and partially reversed following the public release of ChatGPT in late 2022. This reversal is most pronounced in fields where coauthors' contributions are more easily replaced by AI, and is observed among both established researchers and newcomers who previously only coauthored papers. The study suggests that generative AI is enabling more researchers to publish solo work, particularly in computationally oriented topics.

Why it matters: This research provides empirical evidence that generative AI is reshaping scientific collaboration by substituting for human labor, altering the traditional division of cognitive work in research.

ResearchOfficialarXiv Computers and Society

LLMs Fabricate Legal Citations at High Rates for Saudi Law but Not GDPR, Study Finds

A new bilingual benchmark study shows that freely accessible large language models (LLMs) fabricate legal citations for Saudi data protection law (PDPL) in 60-77% of cases, while achieving near-perfect accuracy (94-100%) on the EU's GDPR. The research tested 120 questions in both Arabic and English across three models, revealing that fabrication rates are driven by the jurisdiction of the law, not the language of the query. The study also found that high model confidence does not prevent fabricated citations, highlighting a significant reliability gap.

Why it matters: This research highlights a critical jurisdiction-based reliability gap in LLM-generated legal citations, raising concerns for regulatory compliance and legal decision-making.

ResearchOfficialarXiv Computers and Society

Gauntlet: Multi-Agent LLM Pipeline Outperforms Humans in Deep Technical Critique of Computer Architecture Papers

Researchers present Gauntlet, an open-source pipeline that uses five independent expert-persona LLM reviewers and an adversarial synthesis stage to analyze computer architecture papers. In evaluations on 20 ISCA and HPCA papers, human judges preferred Gauntlet's analyses over those by human experts in 15 out of 20 cases, with statistically significant advantages in critical rigor. Ablation studies show that the multi-agent structure, especially the synthesis stage, is key to Gauntlet's performance gains over single-agent LLM baselines.

Why it matters: This work suggests that structured multi-agent LLM pipelines can exceed human experts in deep technical critique, indicating potential new roles for AI in peer review and research evaluation.

ResearchOfficialarXiv Computers and Society

Commenting with Copilot: A Taxonomy and Multi-Year Analysis of Student Code-Generation Specifications

A four-year analysis of undergraduate programming submissions examines how students use natural-language comments to guide AI code generation. The study introduces a taxonomy covering comment type, code expression level, and code construct, and finds that students primarily write 'What' comments but shift to 'How' comments for procedural tasks. Students tend to focus more on verifying generated code than on revising their comments.

Why it matters: This research offers new empirical insights into how students interact with AI code assistants, informing the evolving role of natural language specifications in programming education.

ResearchOfficialarXiv Computers and Society

DeepBias: Adaptive Framework for In-Depth Probing of Social Biases in LVLMs

Researchers have introduced DeepBias, an adaptive framework designed to probe social biases in Large Vision-Language Models (LVLMs) more deeply than traditional static datasets allow. DeepBias uses a dynamic loop involving a ProposerAgent that generates test data and a DiggerAgent that iteratively rewrites these tests based on model responses, enabling the exposure of progressively deeper biases. The team also developed DeepBiasBench, a benchmark constructed using an ensemble of five state-of-the-art LVLMs to identify vulnerabilities shared across different architectures.

Why it matters: This work advances LVLM safety assessment by introducing an adaptive, evolutionary approach that reveals deeper and more nuanced model biases than static datasets can uncover.

ResearchOfficialarXiv Computers and Society

AI Agent Architecture Strongly Influences Journalism Task Performance, Study Finds

A new preprint systematically compares four AI agent architectures—monolithic, chain-based, multi-agent, and iterative—across 50 journalism tasks using the same language model and tools. The study finds that architecture explains 82% of the variance in processing behavior. Multi-agent collaboration achieved the highest accuracy (84.7%) but required about twice the time of other designs, while the monolithic architecture exhibited a 71.7% source rejection rate, paralleling classic human gatekeeping.

Why it matters: The findings provide evidence-based guidance for newsrooms on selecting AI architectures based on priorities such as speed, accuracy, or auditability.