Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

220 published research events 4 new research papers today

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Google researchers propose a method against test memorization in AI agents

Research from Google Cloud AI Research addresses the problem where AI agents that optimize their own harness overfit too much to tests. The new RRSI method regulates optimization with a budget of edit changes and a filter against task-specific tricks. Result: generalization is preserved without increasing compute costs.

The Decoder (daily AI news) Original source ↗
Research confirmed by 2 independent sources

NASA and IBM released an open-source foundation model for lunar science

NASA and IBM released the Lunar Foundation Model, an open-source foundation model for lunar science. Trained on 2 million tile bundles from 17 years of data from the Lunar Reconnaissance Orbiter and other missions. The model specializes in ice and crater detection.

Part of an overview of multiple AI topics; only this event has been covered.

Root.cz · and 1 more source Original source ↗
Research only one source so far

Chinese AI models are subject to state doctrine, study by Aleph Alpha finds

Aleph Alpha created a benchmark covering 967 politically sensitive topics. Tests of the Qwen, DeepSeek, and Kimi models showed that only 17–41% of responses were balanced; the rest repeat state doctrine, evade the question, or refuse to answer. The bias also shows up in unrelated questions. Nvidia Nemotron contains approx…

The Decoder (daily AI news) Original source ↗
Research only one source so far

According to the report, Anthropic is consulting religious experts on the possible consciousness of Claude models

According to a New York Times report, Anthropic has since autumn 2025 invited dozens of theologians and philosophers to discussions about the possible consciousness and moral behavior of Claude models. The report is based on interviews with 20 participants; this does not prove consciousness of the models.

The Decoder (daily AI news) Original source ↗
Research only one source so far

Study presents Mem++ for preserving document history in LLM agent memory

The study presents Mem++, which stores entire documents with date and author without a generative model at write time. When queried, it selects time-matching documents. According to the authors, it outperformed the strongest comparison memory system on OrgMemBench by 8.0 to 13.1 points.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

SELF-POT study measures the benefit and cost of test-time computation

The SELF-POT study compares five low-cost models across 350 tasks and counts all calls toward the cost. In repeated evaluation of coding candidates, selection based on public examples increased the number of correct solutions from 376 to 453 out of 500 and reduced API costs by 12–49%.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

DSB-DG benchmark reveals loss of document fidelity in voice agents

The study presents the DSB-DG benchmark with 1 636 question-answer pairs from 50 documents across five professional domains. According to the authors, the accuracy of voice agents decreases with context and dialogue length; the ASR-LLM-TTS architecture performs best.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Study introduces CTWM for more efficient memory handling in language agents

The study proposes CTWM, which allocates the context budget according to memory usage frequency. The authors report a token savings of 24.48% on LongMemEval with comparable overall accuracy, and on the Synthetic Graph World a 13.6% decrease in prediction error for the less frequently used half of memory.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Study links the Qwen3.6-35B-A3B model with geometric tools for spatial planning

The authors tested an agent using the Qwen3.6-35B-A3B model with a library of geometric trajectories in a dynamic 2D environment. According to the study, the configuration without reasoning achieved the same goal-reaching success rate as the chain-of-thought variant, while cutting decision time from minutes to seconds.

arXiv cs.AI (Artificial Intelligence) Original source ↗