Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

716 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Graph Theory Agent and a benchmark for graph reasoning in LLMs

Researchers from arXiv introduced Graph Theory Bench with 100 000+ examples (24 graph problems in four representations) and Graph Theory Agent (GTA) to improve model performance. GTA improved Phi-4 from 53.5 % to 69.1 % on easy tasks and from 33 % to 41.5 % on hard tasks. The code is freely available.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Knowledge graph editing and the displacement of correct answers: a study of model locality on FB15k-237

The research examines how edits to knowledge graph models (KGE) can push correct answers out of the top results, even though the edit appears successful. On the FB15k-237 dataset with the DistMult and ComplEx models, direct editing works without harm in only 23 % of cases. Support-regularized editing achieves 36–38 % harmless…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Benchmark Radar: a database and search engine for AI benchmarks

The research team introduces Benchmark Radar – a live database and search engine for AI benchmarks and evaluations. The system catalogs 1283 source records and 12916 observations from 37 data sources, including 13 direct connectors and 24 research feeds. It provides a web dashboard, a leaderboard, a CLI for offline…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Principle-guided improvement of language models from limited text

The research program at Qiushi Engine focused on BabyLM 2026 Strict-Small (10 million words) involved three phases: building a frontier model, discovering principles of efficient learning, and applying them. The result was an increase in the aggregate score from 42.02 to 42.25. The models and code are available on Hugging Face and…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Metrics for evaluating multi-turn AI agents: AEM breaks down quality to the level of individual turns

AWS Machine Learning Blog introduces Agent Evaluation Metric (AEM), a methodology for measuring the quality of multi-turn AI agents. Unlike holistic scores that mask errors, AEM makes it possible to distinguish between an error in a single turn and its cascading impact on subsequent turns. The metric decomposes…

AWS Machine Learning Blog Original source ↗
Research only one source so far

Research: AI agents increase the number of requests in public services

Chris Schmitz presents a study mapping “agentic flooding” – an AI-driven rise in complaints and requests in public services. Analysis of 84 cases in 11 countries: UK ombudsman +170% (2022–2024), CFPB USA +500%, similar trends in Brazil and Germany. The author sees an opportunity to reform services, not a threat –…

TechCrunch AI Original source ↗
Research only one source so far

A framework for geometric reasoning in language models combining diagram analysis and logical deduction

The research team introduces a framework with Geometric Vision Parser and Symbolic Solver modules that improves geometric reasoning in LLMs without the computational overhead of multimodal models. It achieves performance comparable to the Gemini 2.5 Pro model on a benchmark of tasks from Chinese Zhongkao exams in 2025.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Adaptive coupled game modules in artificial general intelligence

A preprint from arXiv introduces a probability wave framework for modeling the behavior of interacting adaptive agents. Using data from the Chinese stock market, it empirically verifies that adaptive coupled models explain 82–94 % of observed behavior (vs. <5 % for purely independent ones), and proposes their…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

MatBrain: a collaborative agent with two models for research into crystalline materials

MatBrain has been published, a collaborative agent with two specialized models (Mat-R1 30B for analytical reasoning, Mat-T1 14B for tool orchestration). The system is lightweight and can be deployed locally, competitive with frontier LLMs. In a catalyst design test, it generated 30 000 candidate structures and…

Nature Machine Intelligence Original source ↗
Research only one source so far

The ERASE research method enables data unlearning during inference without retraining the model

A research paper introduced the ERASE (Erasure via Reconstructive Adversarial Signal Editing) method for selectively forgetting designated data during inference without modifying model weights or retraining. The method uses structured, class-conditioned input perturbations and maintains performance on other…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

DAREBench: a benchmark for evaluating LLM models in agent applications

Researchers on arXiv introduced DAREBench, a benchmark for evaluating LLM models as agents. It contains 233 tasks from 22 sources organized into a matrix by input modality and execution mode. They tested 23 commercial API models and 12 open-weight models across 7587 runs. No model won in every group…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

How reasoning representations support human evaluation of LLMs

The study compared six formats for representing reasoning in LLMs. It found a discrepancy: users prefer representations based on planning and decomposition, but simpler chain-of-thought better supports verification, trust calibration, and interpretability.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

PetQA: A Benchmark for Evaluating Veterinary Knowledge and Clinical Decision-Making in LLM Models

The new Korean benchmark PetQA for evaluating veterinary knowledge in LLM and LVLM models contains 10 076 text-based and 8 751 multimodal QA pairs provided by veterinarians. The study evaluates 18 models using ROUGE, BERTScore and LLM-as-judge metrics with three methods: zero-shot inference, retrieval-augmented generation and…

arXiv cs.CL (Computation and Language / NLP) Original source ↗