Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

825 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

LLMRouter: infrastructure for developing and deploying intelligent LLM routers

A research project presents a unified formulation of LLM routing (5 components: context encoders, model encoders, scoring functions, decision rules, learning signals) and the xRouteBench benchmark for various tasks. The open-source LLMRouter infrastructure includes 16+ routers. Experimental results show…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

FutureBridge: Optimized token selection in collaborative decoding

The research introduces FutureBridge, a technique for improving collaboration between large and small language models. Instead of relying on the local preferences of the large model (LLM), it uses supervised token reranking based on how well the tokens enable the small model (SLM) to continue. On five benchmarks of mathematical…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

An agentic hybrid approach to knowledge graph generation

The paper proposes an agentic approach to generating knowledge graphs for HR platforms. It combines LLMs with Wikidata and processes multilingual skills (5 European languages) in five stages: reconciliation, canonicalization, curation, deduplication and recovery. The system maps unstructured text to structured…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

A benchmark for analyzing LLM agent trajectories with a focus on responsibility attribution

The research paper presents a benchmark and framework for analyzing LLM agent trajectories, focused on attributing responsibility for individual steps. It includes 1300+ annotated trajectories from AgentDojo and Agent3Sigma, defines two evaluation tasks and provides baseline results. It releases reusable annotation…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Agent decision-making strategies under uncertain perception

The study compares agent behavior in the Artificial Life model under noisy perception. Agents that are aware of uncertainty have a significantly higher survival rate and fewer critical errors. The results show a qualitative shift from exploratory to cautious behavior as uncertainty increases; explicit collection…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Diagnosing failures in multimodal speech models: separating decision-rule flaws from representation readout limitations

A research team proposes a diagnostic method that separates different sources of failure in multimodal models for speech and emotion analysis. The method compares the emitted answer, option logits and different ways of reading out the hidden state. Across five systems and two emotion datasets, an average improvement was found…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

SEE benchmark: multimodal LLMs cannot reliably reason from experimental data

The new Science Edge Evaluation benchmark tested 19 multimodal LLMs on scientific tasks in chemistry, biology and materials science. The best model achieved 48.7% accuracy, and general-purpose models outperformed specialized ones. Adding tools increased accuracy to 52.7%. The study found that models cannot scientifically…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Tone-Aware RAG: awareness of communication style as a design principle for retrieval-augmented generation

The paper finds that RAG systems ignore user instructions for communication tone (e.g. formal, friendly, simple) because the style of retrieved documents dominates generation. It proposes the TA-RAG framework with four constraints: dishonorable language, readability, adaptation to the recipient, and empathetic framing.…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

New NS-RIS algorithm enables quantum Markov models to outperform classical methods

Researchers introduced the NS-RIS (Newton-Schulz Retraction-based Inference) algorithm for more efficient training of Hidden Quantum Markov Models (HQMMs). NS-RIS demonstrates in practice for the first time that HQMM can outperform classical EM-trained HMM even on data not generated by quantum processes. On synthetic HMM…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Generative AI in mathematics: new discoveries and ethical questions

OpenAI announced ten mathematical discoveries made with the unreleased Astra model (geometry, cryptography, coding theory). The mathematics community is discussing questions of attribution and AI's impact on the discipline's future. The Leiden Declaration advocates AI as support for human creativity rather than a replacement; leading mathematicians…

The Conversation — Artificial Intelligence Original source ↗
Research only one source so far

Readers rate ChatGPT 4.0-generated short stories more highly than human writing when they do not know their origin

A study with 2500+ participants showed that readers cannot distinguish ChatGPT 4.0-generated short stories from human-written ones better than chance and rate AI texts more highly (quality: 1.54 vs 0.97; immersion: 1.42 vs 1.00). However, ratings fall once readers learn that the texts were generated by AI.

The Decoder (daily AI news) Original source ↗
Research only one source so far

WeatherNext outperformed existing forecasting techniques and gave meteorologists an extra day to issue warnings

WeatherNext, developed by Google DeepMind and Google Research, can forecast hurricanes more accurately than existing models. According to research published in Nature, it provides an average of one additional day of warning compared with existing models — its three-day forecasts are as accurate as…

Ars Technica (AI) Original source ↗
Research only one source so far

AI agents consume 600 times more energy than simple chat

Climate scientist Zeke Hausfather analyzed eight weeks of Claude Code usage: 1 138 prompts triggered over 14 000 model calls and processed 3.2 billion tokens, consuming 170 kWh (150 Wh per prompt). This is roughly 600 times more than Google and OpenAI figures suggest.

The Decoder (daily AI news) Original source ↗
Research only one source so far

TutorMoments: Measuring LLM tutors' ability to distinguish when to help and when to let students work

Hugging Face introduced TutorMoments, an evaluation framework based on real tutoring sessions that measures LLM models' ability to choose between support and independence. The result: Models provide too much help; an explicit instruction improves performance, but a gap remains compared with human tutors. It releases the dataset, code and…

Hugging Face Blog Original source ↗
Research only one source so far

Scientists create 16 new viruses using AI

Scientists from Stanford University and Arc Institute used Evo 1 and Evo 2 AI models to design 16 entirely new, functional bacteriophages based on genomes from millions of organisms. Of 300 synthesized candidates, 16 proved fully functional.

Wired — AI section Original source ↗
Research only one source so far

Detecting AI disinformation through analysis of discussion manipulation

The research team analyzed over 1600 comments under BBC News videos on YouTube and found that a more effective way to detect AI disinformation is to recognize red herring tactics and diversion of conversations toward polarizing topics rather than traditional AI-written text detection, which generative AI…

The Conversation — Artificial Intelligence Original source ↗
Research only one source so far

Employers value soft skills as AI automates entry-level positions

A survey of 647 employers showed that communication, teamwork and critical thinking are valued more highly in graduates than AI knowledge. Automation eliminates traditional entry-level roles such as apprenticeships; higher education still prepares students for the old model.

The Conversation — Artificial Intelligence Original source ↗
Research only one source so far

PoolBench: a benchmark for evaluating pooling strategies for concept representation in decoder-only models

The new PoolBench benchmark evaluated 19 pooling strategies on 3 open language models (Llama, Gemma, Mistral) using 37 693 texts. The W4_hierarchical strategy achieved an AUROC of 0.7799, significantly better than the standard P1_last_token method (0.7640, p=2.0e-36). The research revealed that the choice of construction method has…

arXiv cs.CL (Computation and Language / NLP) Original source ↗