Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

419 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Test-time scaling: candidate generation strategy affects LLM energy consumption and latency

The study finds that test-time scaling for LLMs consumes significantly more energy when candidates are generated sequentially rather than in a batch. Eight sequential calls on an A100 GPU consume 4.64–4.86x more energy and have 5.77–6.12x higher latency. Measured using Phi-3-mini and Qwen2.5-1.5B on the GSM8K dataset.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

PosteriorBench: evaluating the posterior properties of generative inverse solvers

The new benchmark PosteriorBench evaluates whether generative models solving scientific inverse problems actually capture the correct posterior distribution. The benchmark includes four physics problems and five metrics (maximum mean discrepancy, Wasserstein distance). The results reveal gaps in posterior calibration…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Chain-of-thought entropy as a reliability signal: a preregistered replication

A replication study confirms that the shape of entropy in the computational chain of a language model predicts answer correctness on mathematical benchmarks. Tested on four open-weight models on GSM8K and MATH-500, with accuracy differences of +9.6 to +27.5 percentage points. The original claim about the magnitude of the drop…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

VisKG-LM: Visual Memory of Knowledge Graphs for Question-Answering Systems

The research introduces VisKG-LM, a method that compiles knowledge graphs into images offline and caches them as visual memory. A language model then uses this memory without the need for repeated encoding. On the CommonsenseQA benchmark, it improved results by 1.2 points, on OpenBookQA by 0.8 points, and on MedQA-USMLE by…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

BioPhys-Bridge benchmark for evaluating the scientific reasoning of language models in biophysics

A new benchmark for evaluating language models in biophysical research, with 500 cases and 1517 tasks across 6 biological domains. The model DeepSeek-V4-Flash achieved the highest F1 score of 0.360, followed by Qwen3.7-Max (0.316) and GPT-4o-mini (0.294). The dataset is freely available on GitHub and Hugging Face.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Causal mechanisms of hidden information transmission in language models

A study on arXiv examines how Llama-3.1 (8B, 70B) and Qwen models transmit hidden information. Causal analysis shows that internal states at certain network depths control hidden information better than the geometry of output vectors. The study also reveals multi-token confusion.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Regularized emphatic temporal-difference learning: Stability with constant step sizes

The research paper introduces the RETD algorithm, which addresses stability issues in the emphatic temporal-difference learning algorithm in off-policy settings. The paper includes theoretical convergence proofs for both harmonically decreasing and constant step sizes, along with experimental validation on several tasks.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

ReaLMem: A Long-Term Multimodal Memory Benchmark from Personal Archives

The research team published ReaLMem, a long-term memory benchmark built from authentic personal photo archives. It evaluates models at three levels: factual recall, personality inference and predictive personalization. The team proposed ChronoProfiler with temporal weighting to resolve conflicts between preferences. Tests…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Interoception and cybernetics in autonomous AI agents

A framework combining cybernetics with interoceptive inputs enables autonomous agents to monitor their internal states and autonomously update their goals. It brings together principles from artificial intelligence, robotics, neuroscience, and cognitive science.

Nature Machine Intelligence Original source ↗
Research only one source so far

Adaptive activation steering in generative models

Research from Apple Machine Learning Research introduced the DSAS method, which selectively modulates the strength of interventions that steer the behavior of generative models. The method improves the trade-off between mitigating toxicity and preserving quality without significant computational overhead and is applicable to both LLMs and diffusion models.

Apple Machine Learning Research Original source ↗
Research only one source so far

Pew Research survey: Concern about the impact of AI on employment prevails worldwide

Pew Research surveyed 42 151 people from 37 countries (February–May 2026). In 34 of the 37 countries, the prevailing belief is that AI will lead to job losses. Concern is highest in the USA (71 %), Australia and South Korea (76 %). 41 % of respondents feel a mix of excitement and concern, 37 % are primarily concerned, 13 % are primarily…

Research only one source so far

MIRAGE: How conversation state affects the use of historical information in multimodal personal agents

The MIRAGE research study evaluates how multimodal agents use historical information in long conversations. It found that agents generate plausible responses even without access to the history. Open-source models are heavily dependent on context continuity and do not automatically opt for alternative retrieval.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Energy efficiency of AI agents in edge-cloud infrastructure

A study on arXiv introduces the agentic-eCAL metric for evaluating the energy efficiency of multi-agent AI workflows. Experiments on GPUs from NVIDIA (A100, H100) with 16 models show that transferring text between agents accounts for 0.25 % of the total energy, while the dominant costs stem from additional processing…

arXiv cs.AI (Artificial Intelligence) Original source ↗