Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

825 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

EnvACE: Training LLM agents without external environments through internal modeling

The research team introduced EnvACE, a method replacing costly interactions with external environments during training with ‘world rehearsal’: The agent alternates between generating actions and playing the environment's role. It achieves strong performance on BFCL-v4, tau²-Bench, VitaBench and FinMCP-Bench. The source code is…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Agentic AI in integrated sensing and communication

The research survey introduces AISAC, a paradigm for AI agents in sensing and communication. The framework contains six phases (observation, contextualization, reasoning, planning, execution, feedback) and five maturity levels. An audit of systems found that none reports more than 1–2 of the nine agentic capability criteria.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

SkillTrace: Auditing skill reuse in LLM agent ecosystems

The SkillTrace research framework audits skill reuse in LLM agent ecosystems. It extracts three provenance traces (Expression, Implementation, Operational) and represents the flow as a Skill Operational Graph. It achieves AUROC 0.938 and F1 0.898 on 820 transformed positives. An audit of 36 446…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

MoCA: A benchmark for analyzing implicit social context in multimodal data

Scientists introduced Implicit Social Context Analysis (MoCA), a task with a dataset benchmark containing 3 108 multimodal instances for studying affect, intent and stance in human communication. Tests showed that current multimodal language models struggle with this task; a framework was proposed…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Unified Agent: Managing interactions across devices

An arXiv paper on the Unified Agent system for AI agents working across multiple devices over time. It proposes efficient state management for cross-device interactions. Measurements show that it significantly outperforms four comparison approaches and remains robust across different multimodal LLM models. Code and data will be on GitHub.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Conditioned cognitive biases in LLMs: how biased conversation affects in-context reasoning

The study examines how biased conversational interaction in a multi-turn setting affects the expression of cognitive biases in LLMs. It tested eight models on 24 300 jury-validated prompts covering 81 cells of a 9×9 matrix. It found that biased conversation increases the expression of biases in six out of eight…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Analyzing number and date localization in LLMs

The research tests five LLMs on localization of numbers, times and dates. Including localization principles in the prompt produced a statistically significant improvement in accuracy compared with direct translation and alternative strategies.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Marginal matching does not induce factorization in generative models

The study proved that marginal matching (a common regularization approach) in factorized generative models does not prevent the latent style variable from containing class information. Although the model achieves near-zero global MMD, a linear probe still recovers the class with 74–100% accuracy (compared with 10%…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Comparing diffusion and autoregressive language model performance

The study compares diffusion language models (DLM) and autoregressive models (ARM). DLM achieve higher arithmetic intensity through parallelism but do not scale efficiently to long contexts. Blockwise decoding improves scaling. ARM show better throughput in batch inference. The key to lower DLM latency is reducing sampling…

Apple Machine Learning Research Original source ↗
Research only one source so far

ARBITRAGE: more efficient speculative decoding for model reasoning

Researchers from UC Berkeley and Lawrence Berkeley National Lab proposed ARBITRAGE, a framework for step-level speculative decoding. Instead of a fixed acceptance threshold, it uses a router trained to predict when the target model produces a better step. On mathematical benchmarks, they achieved inference speedups of up to 2× with…

Apple Machine Learning Research Original source ↗
Research only one source so far

Female-dominated occupations are most exposed to AI automation

An Australian study shows that administrative occupations (70% women) are most exposed to AI automation. Of the 20 most at-risk occupations, 15 are female-dominated; among the least at-risk, 17 are male-dominated. Vacancies in exposed occupations are falling by up to 22% annually.

The Conversation — Artificial Intelligence Original source ↗
Research only one source so far

Genomic models design new viruses that infect bacteria

Stanford University scientists used large-scale genomic models to design new viruses infecting bacteria. All created viruses are closely related to existing ones and have some different properties. The researchers highlight the possibility that similar AI could in future design viruses targeting…

Ars Technica (AI) Original source ↗
Research only one source so far

WeatherNext improves hurricane forecasting by a day

Google DeepMind's WeatherNext improves hurricane forecast accuracy by an average of one day. Three-day forecasts are as accurate as previous two-day forecasts. Trained on a combination of general meteorological and hurricane-specific data, the model handles both track and intensity prediction.

Wired — AI section Original source ↗
Research only one source so far

Research reveals factual errors in AI chatbot advice

British organization Which? tested AI chatbots on financial advice. The tools correctly explained general concepts but made errors in data — ChatGPT cited nonexistent products, while Copilot used incorrect rates. AI drew from social networks and did not provide current regional information.

Research only one source so far

MCTS-Report: An algorithm for automated report generation with charts from tabular data

The research team proposed MCTS-Report, a framework combining Monte Carlo Tree Search with language models to automatically create professional reports from structured data. It divides generation into atomic actions (chapter planning, chart creation, formulating insights), which it optimizes according to…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Grounding visualized questions: A weakness of multimodal LLMs and a new improvement method

The study finds that multimodal language models perform an average of 17.8 points worse on tasks when the question is presented as text in an image rather than plain text. Although models transcribe the text correctly, they cannot use it as an instruction for reasoning. The authors propose prompt-region grounding, which…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Observation-Calibrated Self-Distillation: a new method for training AI agents with higher accuracy

The research paper introduces the OCSD method for training language models for agents. It addresses the problem of confounding in supervision by using two structurally identical replay views — with and without the future observation. Tested on ALFWorld, WebShop and Search-QA, the method consistently outperformed strong baseline models.

arXiv cs.LG (Machine Learning) Original source ↗