Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

825 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Mitigating rubric interference in LLM judges through on-policy self-distillation

The study identifies rubric interference: when an LLM assesses multiple criteria simultaneously, its verdicts change depending on the combination of criteria (only 1/3 of samples have consistent verdicts). The authors propose SARA, which uses single-rubric assessments as anchors and on-policy self-distillation. Validated on…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Policy algebra for reliable AI agent execution

Researchers proposed a formal framework (policy algebra) for safe execution of AI agents in enterprises. It defines agent reliability as the ability to achieve a goal without violating constraints on data access, authorization, side effects and budgets. Evaluation: it stops 94.8% of policy violations with 86.9%…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

UniFed-VLM — federated instruction tuning of vision-language models

The research paper proposes UniFed-VLM, a method for federated instruction tuning of vision-language models. It addresses heterogeneity in client tasks, modalities and architectures through FedCSA for adapter aggregation and TCoD for knowledge transfer. Code is available on GitHub.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Unwritten benchmark: a new test of the limits of multimodal models in abstract perception

A research team presented the Unwritten benchmark on arXiv to test the ability of multimodal models to recognize words from acoustic signals produced by a pencil and video of its movements without visible ink. While humans achieve accuracy of >80 %, the GPT-4o and Gemini 2.5-Pro models fail with accuracy below 10 %. Paradoxically, the combination…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

New RHMP model for neural simulations of physical fields on meshes

The new RHMP architectural approach simulates physical fields and separates topological aspects (conservation laws) from geometry learned from data. Tested on 7 physical domains, including fluids, electromagnetism, and CFD. It achieves the best performance particularly in the interplay between topology, geometry, and field structure.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Measuring AI efficiency: a replication study questions the α-FLOPs formula on newer hardware

A replication study verified the α-FLOPs formula for predicting AI operation execution time. It confirmed that raw FLOPs are not a sufficient metric but found that the formula falls short on newer hardware — with oscillations and abrupt changes in execution time that it predicts inadequately.…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

AI-designed bacteriophages and the regulatory challenges of technological convergence

Scientists created AI-designed bacteriophages capable of killing antibiotic-resistant strains of E. coli. The article points out that biotechnology regulation is not keeping pace with the convergence of AI, CRISPR and synthetic biology, and laws designed for individual technologies cannot address their combinations.

The Conversation — Artificial Intelligence Original source ↗
Research only one source so far

A specialized semismooth Newton method for kernel-based optimal transport

Scientists from Apple ML Research, MIT and UC Berkeley proposed a new semismooth Newton method for solving kernel-based optimal transport problems. It outperforms the computationally expensive SSIPM approach and achieves O(1/√k) global and quadratic local convergence. Experiments show significant speedups…

Apple Machine Learning Research Original source ↗
Research only one source so far

DiG-bench: a new benchmark for discovering environment rules

DiG-bench is a new benchmark of 70 text-based games designed to measure the ability of AI models to discover hidden rules in an environment. A research team from Oxford, Princeton, MIT and other institutions released 21 games publicly, while the remaining 49 are hidden. The games remain unsolvable by current frontier models.

Import AI (Jack Clark, Anthropic co-founder) Original source ↗
Research only one source so far

Axiom Math verified the proof of the 246 theorem on prime numbers

Axiom Math used its AxiomProver AI system to verify the proof of the 246 theorem on prime numbers — a fundamental result in prime number theory. The verification demonstrates the potential of automated proof checking; the system created a library of components for reuse in further research.

IEEE Spectrum — Artificial Intelligence Original source ↗
Research only one source so far

Evaluating agentic learning harnesses without labeled data

The research paper proposes a method for evaluating continuous learning in agent systems without labeled benchmarks. A stronger teacher model provides sparse, correct corrections to a student with a learning harness; improvement relative to the teacher correlates with actual improvement. The method is validated on cybersecurity tasks and more…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

A taxonomy of misunderstanding in AI-mediated communication: 11 failure modes and 8 analytical layers

The research article identifies 11 specific mechanisms that generate and amplify misunderstanding in AI-mediated communication. It consolidates findings from nine disciplines, models 8 analytical layers and provides an evidence matrix with 9 analyzed dialogue cases. It maps where in the communication process…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Federated Prompt Learning: a framework for privacy-preserving training of large language models

The research survey integrates federated learning with large language models for decentralized training without centralizing data. It analyzes communication efficiency, security and privacy in pre-training, fine-tuning and practical applications. It identifies remaining security and robustness challenges.

arXiv cs.LG (Machine Learning) Original source ↗