Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

825 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

BenchMIRT: a method for auditing what LLM benchmarks really measure

Hugging Face introduced BenchMIRT, a method that analyzes LLM benchmarks at the level of individual questions using multidimensional item response theory (MIRT). The methodology identifies which capabilities individual questions actually measure. Trained on 100 models, 16 benchmarks and 34K+ questions, it independently identified two…

Hugging Face Blog Original source ↗
Research only one source so far

OCR-MetaReasoning Benchmark for testing reasoning in multimodal LLMs on text-rich images

Researchers introduced OCR-MetaReasoning Benchmark with 1500 samples for evaluating multimodal LLMs. The benchmark distinguishes three types of logical reasoning (deduction, induction, abduction) and separately evaluates answer correctness and adherence to the process. Tests show that models struggle with inference that is sensitive…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

ReVA: Research into a region-aware multimodal model for questions about images

The research paper introduces ReVA, a model combining CLIP ViT-L/14 and Qwen2.5-7B-Instruct, with a dual-bridge architecture that aligns representations at both the image and individual region levels. It achieves 82.85 % F1 on the POPE benchmark, thereby reducing object and spatial hallucinations in responses to…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

ImageEval 2026: evaluating Arabic multimodal models with a focus on cultural aspects

The ImageEval 2026 research challenge was introduced with two tasks — AynVQA (spoken visual question answering and hallucination detection in English and Modern Standard Arabic) and CRAI-Bench (evaluation of the cultural accuracy of text-to-image generation). 14 teams participated, using zero-shot prompting…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Deploying a legal RAG system for Uzbek with a fine-tuned retriever

The research paper describes a RAG system for legal questions in Uzbek intended for both cloud and on-premises deployment. The team created two benchmarks (178 expert-annotated queries, 504 QA pairs with LLM scoring), released the UTE-1 embedder for Uzbek, and found that the difference between open-source and proprietary models is…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Emotional context increases the tendency of LLMs to support questionable decisions

The research team tested 6 commercial LLMs (OpenAI, Anthropic, Google) on 324 conversations. It found that when a user expresses emotion, the models are significantly more supportive of premature decisions (from 18.6 to 31.5 points). Five models responded significantly; only Claude Opus remained stable. The effect is not caused by…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

WeAgent-MMSearch: multimodal agents with native text-image interaction

The research paper describes WeAgent-MMSearch, a multimodal agentic search system with native text-image interaction. It includes WeAgent-Harness for image persistence and FA-GSPO for error recovery. Agentic post-training improved the average score by 19.22 points, and the model competes with systems with 10x…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Benchmark XHotpotQA for multilingual knowledge composition in question-answering systems

The research team introduced the benchmark XHotpotQA with 15 661 training and 7 405 validation instances with explicit language annotations. The dataset tests the ability of AI systems to work with questions and evidence in different languages; complete language mismatch causes deficits of 10–23 F1 points in answer quality.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Effectiveness of IoT and deep learning for termite detection and severity assessment in tea plantations

Study preprint from arXiv: A framework with Raspberry Pi and a microphone records sounds from tea stems, and CNN networks using spectrograms achieve 81.5 % accuracy in detecting Postelectrotermes militaris termites. The model also quantifies infestation severity by combining probability, amplitude and nearby positions…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Personalized skill routing in LLM agents

The study introduces SkillFeed, a new approach to skill routing for LLM agents that takes the user profile into account. It achieves 75.1 % accuracy, with a 23.1-point improvement over the baseline. On queries with a changed profile, it achieves a 35.1-point gain.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

WikiSkill: a framework for AI agents with persistent memory of errors

Google Research introduced WikiSkill — a framework that gives AI agents a persistent knowledge base of errors and successes. Agents do not learn continuously, but create better instructions (Agent Skills) based on experience. The three-layer architecture (Raw, Wiki, Skill layer) enables gradual improvement…

The Decoder (daily AI news) Original source ↗
Research only one source so far

Accelerating label spreading using algebraic multigrid

The research team proposed AMELS, a method combining faster graph construction algorithms with algebraic multigrid solvers. It achieves a significant reduction in runtime, is more robust to hyperparameters, and enables label spreading on large datasets with a small number of labeled samples.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Study: GPT models for automating systematic literature reviews

A scientific study examined a pipeline for automating information extraction from 536 papers on disease modeling. The GPT-4.1 model achieved 77.95 % accuracy, while GPT-5.0 achieved 81.67 %. Field-level accuracy ranged from 32 to 100 %. The study shows that agreement between models indicates output quality.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Models’ own uncertainty signals rival labeled data when learning to abstain

Researchers found that fine-tuning models (LoRA) based on their own confidence scores without any labeled data is competitive with supervised methods. Tested on 6 open-weights models (1B–8B). The only weakness: the model cannot recognize “confidently wrong facts” – claims it makes confidently…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

GOLLuM: language models calibrated for experimental optimization

The paper presents GOLLuM, a method for training language models using Bayesian objectives. It achieves 40% more efficient experiment search in organic synthesis, materials science, and molecular design; it doubles the discovery of high-performing Buchwald-Hartwig reactions (43 % vs. 24–25 %).

Nature Machine Intelligence Original source ↗
Research only one source so far

Molecular foundation model transfers knowledge across odor recognition tasks

Research shows that the foundation model Uni-Mol2, adapted to predict odor descriptors, successfully transfers to other tasks: odor recognition in individual datasets, binary classification, distinguishing enantiomers, and odor discrimination in mixtures. The model matches or surpasses the previous SOTA on GS-LF…

arXiv cs.LG (Machine Learning) Original source ↗