Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

825 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

LLM translates driver commands into autonomous vehicle parameters

Researchers at TU Delft developed a system that uses an LLM to convert natural language commands (e.g. “I'm in a hurry, drive faster”) into adjustments to autonomous vehicle parameters. The system adjusts the motion planning algorithm with safety in mind, keeps the user involved in the decision-making process, and in simulations…

IEEE Spectrum — Artificial Intelligence Original source ↗
Research only one source so far

A weighted tree structure for memory in long-term agents with large language models

A research paper on arXiv introduces Weighted Memory Tree, a hierarchical memory technique with dynamic retention scores for LLM agents. On the GAIA-Text benchmark with Qwen3-8B, Gemma 4 E4B and Llama-3.1-8B, it achieved an accuracy improvement of 9.97 percentage points and a reduction in token usage of 32.8 % compared with…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

KREL: Automatic Medical Coding with LLM and Knowledge-Based Guidance

The study presents the KREL framework, which combines LLM with external guidelines for ICD coding. The system automatically assigns standard diagnosis and procedure codes to medical reports, reduces LLM hallucinations, and outperforms existing methods on benchmarks.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

AgentMercury: Synthesizing business environments to train AI agents

The research framework AgentMercury synthesizes 4 783 executable environments from business scenarios across 14 industries and 50 countries. Qwen3.5-4B improved performance from 12.3 to 15.7 on EnterpriseOps-GYM and from 45.9 to 56.0 on AIME26. Training the model to construct environments increases the success rate from 3.3% to 83.3%.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

TriPLU: direct trilinear products instead of gated FFN in small language models

The research team introduced TriPLU, a technique using a direct trilinear product of three streams instead of a gated FFN layer. On a 1M-byte prefix of TinyStories, it achieved a validation loss of 1.0637 compared with 1.1017 for SwiGLU. The paper acknowledges that the results are specific to a low learning rate and do not demonstrate general…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

World models of the environment, the agent, and their combined system

The research defines canonical predictive models for three channels: modeling the environment, the agent, and their interactions. It uses computational mechanics and epsilon-machines. It demonstrates how coupling the agent and the environment reduces model complexity. Using a POMDP example, it shows that a constrained model can be finite rather than…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Multiresolution enhancement of implicit neural representations

The research framework WIEN-INR for implicit neural representations distributes modeling across resolution levels and improves the capacity to reconstruct details with a new Enhancement Network. Smaller networks retain the full spatial frequency content with lower requirements for training and scientific data storage.

Nature Machine Intelligence Original source ↗
Research only one source so far

AI for detecting the remaining bruks in the wild in New Zealand

Researchers have developed a method called cross-model confusion mapping that combines microphones and an AI model to detect the last remaining bruks in New Zealand's forests. Instead of training the model directly, they first use BirdNET (a model for identifying 6000+ bird species) to find sounds that AI confuses with bruks, then these…

The Conversation — Artificial Intelligence Original source ↗
Research only one source so far

World models without modeling human beliefs predict the wrong actions, research shows

Research: the new Mental World Modeling (MWM) framework extends world models (Sora, Genie 3, JEPA, Marble) to include modeling human beliefs, emotions and intentions. The authors published the MENTIS implementation, which models both physical and mental aspects of human behavior without additional training in a social…

The Decoder (daily AI news) Original source ↗
Research only one source so far

RecPFN: Prior-Fitted Networks for Context-Based Recommendation

The new RecPFN model brings in-context learning to recommendation. It is trained on synthetic data with a causal prior and predicts the next items from a few examples without updating its weights. It achieves the best zero-shot performance on eight benchmarks and is efficient in deployment.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Automatic recognition of bioinformatics software names in scientific literature

The research paper presents SNAIL, a hybrid system for automatically recognizing the names of bioinformatics software and databases in scientific texts. It combines lexical modeling with transformers (SciBERT) and achieves better results than both bioNerDS2 and general-purpose LLMs (ChatGPT, Gemini, Grok, Claude).

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Model distillation by truncating lower-quality generations: the TUP method without the lower tail

The research team proposes the TUP method, which distills Best-of-N selection (one model instead of selecting from N samples) by removing low-rated generations and increasing the weights of better ones. The method enables offline training with binary cross-entropy without dependence on a specific prompt. The paper includes theoretical justification and…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

MAVEN framework for evaluating multimodal content against macro-societal values

A research team proposed MAVEN, a hierarchical framework for evaluating multimodal content against macro-societal values. The framework organizes values into 6 primary dimensions and 72 indicators and is grounded in international human rights instruments. The authors created the benchmark MacroValue-Bench and a compact…

arXiv cs.CL (Computation and Language / NLP) Original source ↗