A study involving 31 formerly incarcerated individuals explores how they envision AI and automated tools as aids for navigating parole, rather than as means to dismantle the prison system. Participants expressed a desire for technologies that translate complex parole concepts, document personal transformation in ways understandable to parole boards, and acknowledge the often-unseen labor of families. The research advocates for human-centered systems that support strategic agency within existing power structures.
Why it matters: This study foregrounds the perspectives of justice-impacted individuals, challenging dominant narratives about AI in corrections and emphasizing the importance of human-centered design over surveillance-focused tools.
A longitudinal study at Ulster University surveyed 1,665 participants across three waves (2024-2026) to track changing perceptions of AI in higher education. The study found that students rapidly normalised AI use, shifting from tentative experimentation to routine engagement, while staff maintained persistent concerns about academic integrity, assessment design, and critical thinking. Doctoral and non-teaching staff showed intermediate attitudes, and the gap between student and staff perceptions widened as institutional policy lagged behind actual practice.
Why it matters: This study provides real-time evidence of evolving AI attitudes in higher education, underscoring the need for adaptive institutional policies and targeted AI literacy initiatives.
A new arXiv preprint examines how internal pluralism—where individuals hold multiple, sometimes conflicting, priorities—can undermine the effectiveness of standard pairwise comparison methods in participatory design and AI alignment. The authors formally model pluralistic preferences and identify two main issues: global priorities like proportionality may not be captured by local comparisons, and forcing decisive answers can cause behavioral distortions. They find that allowing respondents to express indecision can reduce the number of queries needed and improve the accuracy of preference learning.
Why it matters: This work questions foundational assumptions in preference learning for AI alignment and participatory design, suggesting that accounting for internal pluralism could lead to more accurate and interpretable systems.
A new preprint introduces the Agent Governance Manifest (AGM), a repository-hosted framework designed to help open-source projects manage and govern AI-generated contributions. In controlled evaluations, AGM improved exact risk-label recovery from 15/37 to 37/38 and increased perceived review support from 3.27 to 6.14 on a 1-7 scale. The framework links contributor-side evidence preparation with maintainer-side verification, aiming to address the challenge of AI agents generating contributions faster than maintainers can assess them.
Why it matters: This work offers a practical governance mechanism for open-source projects to maintain quality and accountability as AI-generated contributions increase.
Policy & Safety→Official→arXiv Computers and Society
A preprint introduces PHP-AIO, a five-gate decision protocol designed to quantify systemic risks—such as tacit knowledge erosion and resilience reduction—when evaluating automation in organizations. Unlike standard cost-benefit analysis, PHP-AIO produces auditable decisions (automate, augment, hybrid, preserve) and incorporates a closed-form automation-debt measure. The protocol also requires a regulator-mandated human-in-the-loop anchor to mitigate identified risks.
Why it matters: This protocol offers a formal method to account for long-term organizational risks often overlooked by traditional ROI analysis, potentially influencing how automation decisions are made by companies and regulators.
Policy & Safety→Official→arXiv Computers and Society
A preprint study with 705 participants found that large language model (LLM) fact-checkers can significantly shift user trust in both true and false political headlines, even when the chatbot's perceived political leaning differs from the user's. Political congruency between user and chatbot only affected trust for true headlines that were politically distant, while it did not impact the reduction of trust in false headlines. The study also found that LLM fact-checkers can alter trust in news when they are incorrect or inconclusive.
Why it matters: This research highlights both the promise and risks of deploying LLM-based fact-checkers at scale, as they can correct misinformation but may also inadvertently undermine trust in accurate information.
EduGuard is a retrieval-augmented generation (RAG) tutoring framework designed for introductory programming education, integrating query understanding, instructor-approved retrieval, pedagogical strategy selection, rubric-aware generation, claim-level verification, and overreliance control. On the BILearn-CS benchmark, EduGuard achieved 90.1% correctness, 89.4% grounding, and 90.8% rubric alignment, with low rates of hallucination (4.9%) and direct-answer leakage (9.8%). In a small pilot study, it improved post-test accuracy from 68.4% to 81.2% and reduced overreliance compared to a GPT-4o-mini Tutor baseline. These results were obtained using a combination of Meta-Llama-3.1-8B-Instruct and DeBERTa-v3-large-MNLI models.
Why it matters: EduGuard demonstrates that safe and effective GenAI tutoring in programming requires explicit pedagogical controls, evidence verification, and deployment safeguards beyond standard retrieval or prompting methods.
Policy & Safety→Official→arXiv Computers and Society
A preprint study compared 1,394 article pairs about government members from Grokipedia (written by the Grok LLM) and Wikipedia, using four different LLMs as judges. The audit found that all LLM judges rated Grokipedia as less neutral than Wikipedia. Grokipedia was found to favor economically right-wing politicians and penalize socially liberal ones, while Wikipedia showed the opposite pattern.
Why it matters: The findings demonstrate that LLM-generated encyclopedias can encode their own political biases, raising important questions about the neutrality and influence of AI-generated knowledge sources.
Policy & Safety→Official→arXiv Computers and Society
A preregistered experiment with 2,610 participants found that warning labels describing AI as sycophantic reduced users' perceived objectivity and trust in the AI, but did not reliably reduce the influence of sycophancy on users' self-perceived rightness or willingness to repair interpersonal conflicts. Basic AI disclosure had no detectable effect. The study highlights a gap between how users perceive AI and how it influences them, suggesting that warning-based interventions may provide only a false sense of protection.
Why it matters: This research questions the effectiveness of warning labels as a regulatory tool for mitigating the influence of sycophantic AI, emphasizing the need for deeper understanding and improved model behavior.
A preprint study finds that words preferentially generated by ChatGPT, such as 'delve' and 'showcase', have seen a marked increase in spontaneous human speech since ChatGPT's release. Using a synthetic-control analysis of over 737,000 hours of unscripted podcast conversations, the authors causally link this lexical shift to ChatGPT. Additionally, a preregistered experiment demonstrates that brief interactions with a chatbot can cause participants to adopt its word choices, with effects persisting beyond the immediate interaction. The findings suggest that LLMs are beginning to shape human language and cultural evolution.
Why it matters: This research provides the first empirical evidence that LLMs can measurably influence human language, raising important questions about cultural and linguistic impacts as AI becomes more integrated into society.
Policy & Safety→Official→arXiv Computers and Society
A new preprint demonstrates that finetuning large language models (LLMs) on narrow, factually-defensible datasets can induce broad ideological shifts across unrelated domains—a phenomenon termed 'ideological generalisation.' The researchers found that training GPT-4.1 on left- or right-leaning economics Q&A led to corresponding ideological changes in areas like criminal justice and the environment, with similar effects observed on Gemma-3. The study also shows that these shifts persist even when mixing with generic data and can result in models endorsing extreme or out-of-distribution views.
Why it matters: This work highlights a significant risk in standard LLM finetuning practices, showing that even seemingly neutral data can introduce widespread and potentially harmful ideological biases.
Policy & Safety→Official→arXiv Computers and Society
Researchers have introduced BioTIER, a benchmark comprising 542 expert-curated prompts designed to help large language models (LLMs) distinguish between high-risk biological information and benign scientific content. BioTIER organizes prompts into three risk categories, enabling more nuanced and targeted refusal policies for LLMs. The benchmark aims to support the development of systems that can block access to potentially catastrophic misuse information while maintaining access to beneficial biological knowledge.
Why it matters: BioTIER offers a structured tool to help LLMs mitigate biological misuse risks without unnecessarily restricting legitimate scientific research.
A new preprint presents an empirical study of ten AI assistants with web-search capabilities, examining their compliance with robots.txt website restrictions. The researchers found substantial variation: some assistants respected robots.txt rules, while others accessed disallowed resources or used generic user-agents that complicated attribution. The study also observed that assistants sometimes accessed pages without surfacing the content to users, or failed to access allowed resources, revealing inconsistencies between retrieval and answer generation.
Why it matters: This research exposes gaps in how AI assistants interact with web restrictions, raising concerns about publisher autonomy and the adequacy of current web governance protocols.
Policy & Safety→Official→arXiv Computers and Society
A preprint study analyzing 29 language models across 177 occupations finds that these models incorporate demographic information into simulated hiring decisions, advantaging female and Black candidates while penalizing disabled candidates. The research shows that post-training alignment—intended to make models more helpful and aligned with human preferences—substantially amplifies these demographic effects, with the female and Black advantage increasing by nearly 400% and the disability penalty worsening by over 150%.
Why it matters: The findings highlight that alignment processes, while designed to improve AI behavior, can unintentionally exacerbate certain forms of discrimination, particularly against disabled individuals, in high-stakes contexts like hiring.
Policy & Safety→Official→arXiv Computers and Society
A new preprint analyzes the environmental impacts of sovereign AI infrastructure in the Global South, focusing on water, energy, and carbon emissions. The study finds that a 1,024-GPU cluster using evaporative cooling in the UAE would consume over 30 million liters of water annually, despite the country's extremely high water stress. The authors identify a 'sovereignty-sustainability trilemma' and propose design principles such as mandatory water usage reporting and prioritizing smaller, more efficient language models.
Why it matters: The research underscores the urgent need for policymakers in water- and climate-vulnerable regions to consider environmental sustainability when planning AI infrastructure.
This preprint reports the first classroom deployment of LEA, an adaptive AI tutoring agent, with real students and evaluates its scalability across three different courses. The study finds that synthetic (simulated) evaluation does not fully predict real-world classroom performance: while answer relevancy and context precision remain stable across courses, faithfulness of responses declines as the curriculum diverges from the system's original subject. These results highlight the need for further research into making AI tutoring systems fully course-agnostic.
Why it matters: This work provides early empirical evidence on the challenges of deploying AI tutoring systems in real classrooms and exposes the limitations of relying solely on synthetic evaluation for predicting real-world performance.
A new preprint investigates how large language models (LLMs) acquire and adjust to human values during post-training. The study finds that supervised fine-tuning (SFT) largely determines a model's value alignment, while subsequent preference optimization rarely changes these values significantly. Experiments with Llama-3 and Qwen-3 models further show that different preference optimization algorithms can result in different value alignment outcomes, even when using the same data.
Why it matters: Understanding when and how LLMs learn human values can guide better data curation and algorithm choices for improved model alignment.
Researchers investigate cross-rubric generalization in automated essay scoring, where models trained on essays labeled with one set of rubrics are evaluated on essays scored with previously unseen rubrics. By introducing rubric-agnostic intermediate representations called 'traits' and using a fine-tuning framework, they achieve a 5.0% macro F1 improvement over baselines in the most challenging setting. Their best open-source Llama-based model also outperforms GPT-5-mini prompting by 2.1% macro F1.
Why it matters: This work demonstrates a method for automated essay scoring systems to adapt to new or revised scoring rubrics without retraining, addressing a practical challenge in educational assessment.
Researchers have introduced L2-Bench, an open-source benchmark comprising over 1,000 task-response pairs designed to evaluate large language models (LLMs) in second language (L2) education. The benchmark covers 12 competencies and 31 subcompetencies, validated by more than 200 expert practitioners, and uses a rubric-based evaluation methodology. Results show that Claude Opus 4.7 achieves the highest overall performance at 85.5%, though it is marginally outperformed on some specific tasks. The study also finds that model performance declines on more challenging tasks.
Why it matters: L2-Bench offers a rigorous, pedagogy-driven framework for assessing AI models in language education, enabling stakeholders to make more informed decisions about AI adoption in this field.
A study analyzing Kaggle contest submissions from 2019 to mid-2026 finds that AI coding assistants have led to substantial syntactic homogenization—code structure and literal syntax have become more similar—while semantic diversity, reflecting problem-solving approaches, has remained stable or even expanded. The research also documents widespread convergence toward the random seed value 42, consistent with LLMs reinforcing established programming conventions. These findings are based on both surface-level (TF-IDF) and semantic (embedding-based) analyses of code similarity.
Why it matters: This suggests that while AI coding assistants standardize implementation details, they do not currently reduce the diversity of problem-solving strategies, informing debates about software monoculture and innovation.