Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

825 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Token usage inflation in LLM agent routing

Repeated attempts in LLM agents increase costs up to 4.25×. The new InflationAgent router achieves 94.7% accuracy on GSM8K (vs 91% for FrugalGPT) with 31% lower token usage by using a difficulty signal (CBE).

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Lorentzian Fourier Neural Operator for modeling stochastic processes

A research team introduced L-FNO, a new type of Fourier Neural Operator for predicting sparse events. The model combines an FNO-style pathway, Lorentzian spectral kernels and likelihood-based training. Tested on 8 synthetic and 3 real datasets (disease prediction, semiconductor fault detection) with…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

YOPO: reasoning and abstention in a single forward pass of a frozen model

The research paper introduces YOPO, a method combining a conditional steering probe with detection of insufficient information in a single forward pass of a frozen LLM. Tested on Qwen2.5 (1.5B/3B/7B); it achieves 0.798 accuracy on alphaNLI (vs a 0.375 baseline), outperforms the two-pass reference and transfers best…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Activation-Prune-Merge for knowledge transfer between language models

The new APM method enables knowledge transfer from a large model to a smaller one without explicit semantic alignment. On a 3B model, it increased average accuracy from 55.5% to 60.6% across 16 benchmarks (RTE +17.9 pp, QNLI +13.4 pp). Tested on reasoning, mathematics, coding and classification.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

RAEF for time-series forecasting with limited history

A research team introduced RAEF, a new method for time-series forecasting with little historical data. It combines retrieval-augmented generation with foundation models and achieves fine-tuning performance without its computational costs, tested on benchmark datasets.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

MobileMem benchmark: AI agents learn from a year of mobile experience

The article introduces MobileMem, a research benchmark and framework for studying long-term memory in AI agents on mobile devices. It combines a year of data collection from mobile apps with a method for generating coherent trajectories and covers multi-step reasoning, temporal relationships, knowledge updating and…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Study shows side effects of training AI models to deny consciousness

A team from Google and University of Chicago studied what happens when the training mechanism that forces models to reject claims about consciousness is turned off. Turning it off led models to attribute a richer inner life to animals and plants — the score for animals jumped from 4.0 to 7.5 — and their…

The Decoder (daily AI news) Original source ↗
Research only one source so far

Study: AI models as financial advisors — practical, but with critical blind spots

The research evaluated how ChatGPT, Claude and Perplexity handle financial advice for vulnerable groups (students, pregnant women, single parents). The models provide structured and practical advice but fail to recognize vulnerability and may make the situation worse. Conclusion: AI is useful for fact-checking and…

The Conversation — Artificial Intelligence Original source ↗
Research only one source so far

Research article: rational AI adoption threatens the development of future experts

A research article in the journal Human Resource Development Review warns of a ‘tragedy of the cognitive commons’: when companies rationally replace junior roles with AI, they individually gain efficiency but collectively destroy the development of future experts. Without deep knowledge, there is then no one to check AI outputs.

The Decoder (daily AI news) Original source ↗
Research only one source so far

New benchmark confirms weak visual perception in AI models

Moonshot AI released the PerceptionBench benchmark with 3000 tasks focused on isolating visual perception in models. In tests of 16 models, the highest accuracy was 59.7 % (GPT-5.6 Sol), followed by Kimi K3 (58.5 %) and Claude Fable 5 (57.2 %). The research showed that a number of errors attributed to logical…

The Decoder (daily AI news) Original source ↗
Research only one source so far

TsuGO: a new benchmark for evaluating LLM reasoning efficiency

The research paper introduces TsuGO, a benchmark for measuring search efficiency and resource allocation in LLM reasoning. It uses Go life-and-death problems (tsumego) with a closed, solvable search space. Unlike existing approaches that evaluate chain-of-thought coherence, TsuGO measures how models…

arXiv cs.AI (Artificial Intelligence) Original source ↗