Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

825 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

ViSAGE: self-correcting memories for understanding long videos

A research paper on arXiv introduces ViSAGE, a framework for multimodal agents working with long videos. The system constructs entity-centric memories using cross-modal binding and bidirectional refinement, which minimize entity mix-ups and hallucinations. It achieves 5.9 % higher accuracy than previous…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Architecture of agentic AI systems: layered integration of Ollama and OpenClaw

A research paper describes a layered Agentic AI architecture that separates inference, orchestration, and execution. It integrates Ollama (LLM inference) with OpenClaw (orchestration) into a full-stack system. Validation shows that persistent memory, tool use, and adaptive decision-making emerge from system integration.…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Agents learn predictive knowledge tied to state; the SKL study outperforms trajectory reflection

The research team introduced Stateful Knowledge Learning (SKL): a method that trains agents to extract state-grounded predictive knowledge from experience instead of traditional trajectory-level summarization. Tested on WebShop, ScienceWorld and ChessPuzzles; the results show improvements over reflection-based…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Curriculum training for relational Prior-Data Fitted Networks

The study shows that structured ordering of synthetic data increases the training efficiency of relational Prior-Data Fitted Networks. A single-table curriculum achieved 0.703 ROC-AUC with ~13 300 data points (45× fewer than the baseline), while a relational curriculum achieved 0.638 ROC-AUC with 5 500 data points (220× fewer, 88% of baseline performance). Without…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

LARA: adaptation in residual streams

A research paper on the LARA method for efficient model adaptation has been published. It operates in the residual stream rather than in the weights and matches LoRA in efficiency. It enables smooth control between base and adapted behavior and can run seven adaptations on a single model with 33 MB of overhead.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

LLMs do not understand test difficulty: Limits of automated question generation

An arXiv study examined LLMs' ability to predict difficulty in a large-scale Reading and Writing test. GPT-4 with zero-shot prompting and temperature 0 achieved a QWK of 0.578, but the ConvBERT encoder model performed better at 0.625. LLMs particularly underestimate difficult questions; greater model capabilities lead to even more underestimation.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Deception detection in legal texts: a comparison of transformers and LLMs

The study compares six fine-tuned transformers and seven LLMs for deception detection in legal and general domains. Seven datasets were evaluated (two legal, five general-domain). Sensitivity to domain was found: fine-tuned models lead in data-rich settings, few-shot LLMs are competitive in low-resource legal…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Analysis of the distributed energy requirements of individual chain-of-thought steps in LLMs

The research paper introduces SARE (Step-Aware Reasoning Energy) — a framework for measuring computational effort at the level of individual chain-of-thought steps. The method uses Centered Kernel Alignment between token representations in adjacent layers. Tests on six benchmarks and three open-weight models…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

An LLM framework for discovering mathematical conjectures

The research paper presents a three-stage pipeline for automatically discovering mathematical conjectures: region search from local proofs, reflective validation (groundedness, novelty, potential), and formal verification in Lean 4. Twenty candidates passed through all stages: parsing and type checking in Lean, they were not…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Brain signals improve logical reasoning in language models

The research shows that large language model representations are partially aligned with activity in brain regions related to reasoning. The proposed framework guides LLM representations using fMRI signals, achieving up to a 13% improvement in deductive reasoning accuracy across ten models (1.5–72B…

Nature Machine Intelligence Original source ↗
Research only one source so far

Reinforcement learning for crystal design

Scientists developed a method combining reinforcement learning with generative models for more efficient discovery of new crystals. The approach navigates the candidate space better and enables the design of functional materials beyond the reach of purely generative methods.

Nature Machine Intelligence Original source ↗
Research only one source so far

Quantinuum, Nvidia and Pfizer introduce an AI system for quantum circuit design

Researchers from Quantinuum, Nvidia and Pfizer created an AI system combining transformers and reinforcement learning that designs quantum circuits for molecular simulations a thousand to ten thousand times faster than the existing ADAPT-VQE algorithm. The system generates a circuit in one step and creates solutions with lower…

Part of an overview of multiple AI topics; only this event has been covered.

Research only one source so far

LedgerMind: Multimodal reasoning with provenance verification

The study introduces LedgerMind, a method for multimodal agents ensuring that answers are supported by tools and data. The system normalizes tool outputs into a structured ledger, verifies entity- and number-level grounding and repairs errors as typed states without adding content unsupported by a tool. Tests…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

EvoCause: Improving causal graphs with LLMs for fault diagnosis

The EvoCause study uses LLMs to improve causal graphs for root-cause analysis (RCA) in telecommunications. It presents results on synthetic data (Node F1 higher by 11.59 pp) and the new TeleRCA benchmark with 485 681 alarms. Missing alarm name information reduces performance by 6–8 pp.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

DoTime: A synthetic benchmark generator for interventional and counterfactual causal inference in time series

DoTime was released as an open-source PyPI package with four evaluation suites for testing causal inference in time series. The tool generates synthetic structural causal models with continuous-time interventions, counterfactual sampling, non-stationary dynamics and deterministic profiles (ramp…

arXiv cs.LG (Machine Learning) Original source ↗