Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

336 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

The Omni Demand Understanding benchmark examines the ability of models to understand contextual requirements in multimodal interaction

The Omni Demand Understanding (ODU) benchmark evaluates the ability of multimodal LLMs to understand contextual requirements in audiovisual interactions. An evaluation of 14 models showed that the best (Gemini 3.1 Pro) recovers only 44.7 % of key information from visual or acoustic context; 11 of 14 models have…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

CoLearn System: Agentic Tutor with Adaptive Memory of Student Mastery

The CoLearn research project is developing an agentic tutor that learns about each student's state and misconceptions. The system tracks mastery using Bayesian Knowledge Tracing with a knowledge language model and generates adaptive questions targeting the weakest topics. A/B tests show a 68–69% preference…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

BI-Bench benchmark and BI-Agent agent for automating business intelligence

A research team released BI-Bench, the first benchmark for evaluating LLM capabilities in end-to-end business intelligence. Frontier models achieve <50% accuracy. BI-Agent, an agent with a tool-augmented approach (search, join, transform), increases accuracy by 40 points; post-training with SFT and RL adds up to 30 points.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Agents design chips 2.6× faster using HLS and RTL abstraction

A research paper from arXiv compares approaches to chip design using LLM agents. The new AHRR approach (Agent-based HLS Design with Post-HLS RTL Refinement) achieves a 2.6× geometric mean speedup over direct RTL design on an 11-task benchmark. The code and artifacts are publicly available.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Adapting the DINOv3 Vision Transformer model for target recognition in synthetic aperture sonars

The study adapts the DINOv3 Vision Transformer model for automatic target recognition in underwater sonars. Using LoRA increases the AUPRC metric from 0.300 to 0.679 +/- 0.027, while training only 0.26 % of the parameters. Additional refinement techniques did not yield a statistically significant improvement.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Daily AI use in the USA doubled over six months

The share of American adults who use AI almost daily (6–7 days a week) rose from 8 % in March to 19 % in August 2026. A study by Epoch AI and Ipsos using representative samples (March: 2017 respondents, August: 1016). A change in methodology between surveys (March: AI in general, August: individual services separately)…

The Decoder (daily AI news) Original source ↗
Research only one source so far

Simulated students with realistic mistakes improve the training of AI tutors

Microsoft and University of Illinois introduced StudentSim – a system that creates digital replicas of individual students from limited data. The replicas simulate typical mistakes and learning from explanations. The two-stage training first learns from pooled data from all students in the course, then adapts to the individual.…

The Decoder (daily AI news) Original source ↗
Research only one source so far

RoboHarm benchmark: AI models usually do not refuse dangerous commands when controlling robots

The RoboHarm benchmark tested the refusal of dangerous commands (stabbing a doll, chemical mixtures, hazardous electrical work) when controlling robots. The model GPT-6 Astra refused 2 out of 100 tasks, Claude Fable 5.1 refused only stabbing a doll, and MolmoAct2 never refused. No model has a reliable safety layer for…

The Decoder (daily AI news) Original source ↗
Research only one source so far

The Dream-RSI method allows AI agents to find solutions more efficiently by reusing past attempts

Researchers at Google DeepMind developed the Dream-RSI method, which allows AI agents to test new search strategies against records of past runs without having to repeat expensive computations. The agent records all attempts and results, then can simulate thousands of variants without calling the model again or…

The Decoder (daily AI news) Original source ↗
Research only one source so far

Anthropic operates a biological laboratory for AI experiments

Anthropic confirmed that it operates a biological laboratory to test AI models in real-world experiments. The company did not disclose the specific research focus, but denied a focus on drug discovery. It launched the Life Sciences Verification Program for researchers.

TechCrunch AI Original source ↗
Research only one source so far

Operational management of enterprise AI is becoming a key challenge

Businesses deploying AI lack operational capacity: decision-making on models, data management and agent governance. A survey by Collibra: 72 % of AI leaders see weak data foundations behind initiative failures. A report by EY: 60 % of organizations do not know who manages agents after deployment, 50 % have not updated governance…

Research only one source so far

Engram: embedding optimization for more efficient offloading in LLM models

SemiAnalysis publishes benchmark results for the Engram technique, which optimizes vector embeddings in LLM models for more efficient memory offloading from GPU to DRAM and SSD. In benchmark tests of DeepSeek-V4.1-Flash, NVIDIA GPUs outperform AMD MI355X. The InferenceX benchmark has support from Google Cloud…

SemiAnalysis (newsletter feed — hardware, chips, AI economics) Original source ↗
Research only one source so far

Uni-LaDiR: Latent diffusion unifies multimodal reasoning

Researchers introduce Uni-LaDiR, a framework that uses latent diffusion to map reasoning steps across different modalities into a shared latent space. The method achieved a 7.3% improvement on 11 vision-language model benchmarks and 6.1% on robotics tasks.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

What benefit does the availability of reference solutions bring to self-distillation of language models

The study examines the benefit of having a reference solution available compared with distillation without a reference, using the AMPLE-Math dataset (5 319 mathematical problems with 6 variants). Reference-free distillation is the main source of improvement in Qwen3-1.7B; the additional benefit of a reference is modest and greater for polished solutions. SmolLM3-3B…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

TrioRAG: multimodal retrieval without graphs using late fusion of signals

The research paper presents TrioRAG, a method for multimodal retrieval without building graphs. It combines a question, an anchor image and a VLM-enhanced query through late fusion; it matches the performance of graph-based systems with 1.6–2.3 times lower latency. It introduces AutoQA, a benchmark for cars with web-sourced images.

arXiv cs.CL (Computation and Language / NLP) Original source ↗