Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

825 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

A unified method for detecting hallucinations in multimodal LLMs: the UniHall dataset and SAMF framework

The research paper introduces the UniHall dataset for categorizing hallucinations in multimodal models and the SAMF framework for self-adaptive fuzzing. Experiments reveal performance degradation in state-of-the-art models under stress testing and a trade-off between usefulness and hallucinations in RL-aligned models. Code and…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

OmnilingualGAIA2 reveals a multilingual gap in AI agents

A study using the OmnilingualGAIA2 benchmark in ten languages (five writing systems) tested seven frontier AI agents. All exhibit a cross-lingual gap of 8.8–18.4 pass@3 points concentrated in tool orchestration. The drop does not shrink as model size increases; in non-Latin languages, it is caused by the loss…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

An integrated multimodal AI system for damage assessment with retrieval-augmented generation and sensor fusion

A multimodal AI system for damage assessment combines a locally hosted language model with retrieval-augmented generation, thermal imaging, foundation models and wireless sensing. Knowledge graph-based retrieval outperforms the vector-based approach in cross-document reasoning; multimodal fusion…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

ATLAS: an AI framework for personalized medication safety in older patients with multiple conditions

Researchers introduced ATLAS, a framework that uses LLM agents for medication safety in patients with multiple conditions. The system structures guidelines as a medication safety graph and uses targeted questions to create patient-specific conflict graphs. The GeriMedBench dataset was also created. According to…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

ZeroLock: training language models without backpropagation and with lower memory consumption

A research team proposed ZeroLock, an algorithm that enables large language models to be trained on memory-constrained devices without backpropagation. The prototype reduced memory consumption by 26.5% and increased throughput by 4.9% compared with standard approaches. Its convergence differs from backpropagation by only…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

An adaptive supervised anchor for language model self-distillation

The study addresses weaker supervision in self-distillation when generated sequences deviate from the target. It proposes dual supervision: one for actually visited states and another for correct contexts. Adaptive weighting adjusts to sequence quality. Tests across multiple model sizes confirm improvements…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

OpenVisTool: an open method for synthesizing instructive visual tool-use trajectories

The research paper introduces OpenVisTool, a framework for creating training trajectories for visual tool use in multimodal agents. Key insight: supervision was to be provided only by cases where observations from tools causally contribute to the correct answer. The OpenVisTool-42K dataset contains 42 000…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Researchers propose argumentation as a foundation for explainable AI in decision-making

A research paper proposes computational argumentation as a formal foundation for Evaluative AI – an approach that presents users with competing hypotheses and evidence for and against them instead of a single recommendation. The aim is to create an explicit and contestable system to support human decision-making.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

FineBooks published a benchmark of OCR models to improve training data for language models

The FineBooks project by Hugging Face and EleutherAI tested 14 open-source OCR models on 2165+ historical pages. The best models achieved >97% character accuracy at a cost of <$2 per 1000 pages and are suitable for training LLMs, not for scientific applications. Goal: reprocess 300 000 public-domain books from…

The Decoder (daily AI news) Original source ↗
Research only one source so far

AI for science needs reasoning, not just data

An article in MIT Technology Review argues that AlphaFold is not an ideal model for AI in science. Its success required rare conditions: 170 thousand mapped proteins, 53 years of development and 21 billion dollars. The author therefore proposes AI agents as a better route to accelerating scientific progress.

MIT Technology Review — AI section Original source ↗
Research only one source so far

Transformers are aging: startups develop a new generation of LLM+ models

MIT Technology Review analyzes why transformers — the key LLM architecture since 2017 — are showing limitations. Their dense attention mechanism requires up to 50 million operations for 10 000 words, while OpenAI plans to spend 50 billion dollars on computing resources this year. Transformers also cannot efficiently…

MIT Technology Review — AI section Original source ↗
Research only one source so far

Capek 0.5: a vision-language model for robots with four specialized capabilities

A preprint on Capek 0.5 (a vision-language model for robots, with 2B and 35B-A3B parameters) presents four capability families (spatial reasoning, temporal understanding, action guidance, state verification), trained separately with reinforcement learning and combined through weight-space merging. Tested on benchmarks and in simulation.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Detecting misinformation in the latent space of language models

Research on detecting misinformation through activation engineering in latent space. The method was tested on 11 models (Gemma, Llama, Qwen, 270M–12B parameters) and outperforms the baseline on the LIAR and FACTors benchmarks. It requires neither fine-tuning nor an external knowledge base; code is available.

arXiv cs.LG (Machine Learning) Original source ↗