Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

224 published research events 4 new research papers today

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Continual learning without offline consolidation using isolated replay

The research paper proposes a continual learning method that consolidates knowledge without an offline phase using biologically inspired mechanisms (isolation rule, refractory rotation). On the split-MNIST benchmark, it achieves 91.6 % accuracy, comparable to or better than experience replay, ER-ACE and A-GEM.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Reducing deceptiveness does not guarantee stronger refusal

Research examines whether reducing model deceptiveness via compensatory feature injection improves refusal of harmful requests. On Qwen 3.5 models, they found that a decrease in deceptiveness does not guarantee stronger direct refusal, but under user pressure some versions partially recover (up to 95%).

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Tree Serialization in Language Models: Research on Structure Loss

An Apple Machine Learning Research study tests how well language models serialize tree expressions into natural language and recover them. Results: the channel is asymmetric (up to a 60.4 point difference in accuracy), the best pair achieves 92.9%, 73.6% of failures come from generation. Training on ~3600 examples…

Apple Machine Learning Research Original source ↗
Research only one source so far

Microsoft Research Asia – Singapore continues domain-specific AI research

Microsoft Research Asia – Singapore, opened on 24 July 2025, has published an annual review of its activities in advanced AI research. The institution collaborates with local partners on multimodal healthcare AI, AI agents, and applications for financial services and education. The focus is on translating fundamental research…

Microsoft Research Blog Original source ↗
Research only one source so far

From Research to Practice: Why Deploying AI Vision Outside the Lab Is Challenging

The article explains three main challenges of deploying AI vision from the lab into the real world: models trained on high-quality, controlled data fail under unpredictable conditions; specifically, pose estimation performs excellently in light but fails in darkness, because real-world data with low…

The Conversation — Artificial Intelligence Original source ↗
Research only one source so far

Training world models on video game data

Researchers including Yann LeCun are developing world models — AI models for physical reality — trained on a combination of video and action data. Due to a lack of such data, they are using video game material with millions of games containing control inputs. The startup Worldmodeldata is attempting to be…

Wired — AI section Original source ↗
Research only one source so far

AI algorithm improves thermal stability of RNA vaccines

MIT researchers used an AI algorithm to optimize the lipid nanoparticle formulation in RNA vaccines. The resulting vaccines remain stable at room temperature for up to a year, or at around 38°C for two months. When tested in mice, they elicited the same immune response as a Moderna-like vaccine.

MIT News – Artificial intelligence Original source ↗
Research only one source so far

Study: Cosine Similarity Without Parameters Is Not Proof – Measuring Interpretability Noise in Quantization

A study criticizes the measurement of interpretability during model quantization. Cosine similarity is reported without the parameters (n, ρ) needed for interpretation. On Qwen2.5-1.5B-Instruct: for INT4 the direction rotates beyond noise, for INT8 no change is detected. Released with code and data.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

New method for composing LoRA adapters without interference

A research preprint presents the READ method for composing multiple LoRA adapters into a single model without interference. New adapters can read the inputs of old ones, but not write their outputs. On SuperGLUE, an improvement of more than 20 points; on the domain suite, more than 7 points.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Outdated Documents in RAG: How Old Information Overrides Correct Model Answers

Research reveals that outdated documents in RAG cause errors in 30-37% of Llama and Qwen model responses without instructions; with an explicit instruction, up to 66-75%. The study tests 12 models on a benchmark with 317 verified knowledge reversals in medicine, law, and software. The solution is information about temporal validity.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Self-Play Search Distillation for Improving Reasoning in LLM Models

The paper presents Self-Play Search Distillation (SPSD), a method for generating high-quality synthetic training data through self-play simulations on board games. SPSD leverages MuZero-like networks and produces structured chains of thought. On the Qwen3-4B-Base model, it increases performance in…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

DeepEdu-v1: an efficient local LLM tutor for Vietnamese education

DeepEdu-v1 is an AI tutoring system for Vietnamese education, built on the SCALE framework with optimizations for long context and a self-improving agentic layer. It increases accuracy from 70% to 79.5% on complex tasks, achieves a 2x TTFT speedup compared to vLLM on consumer GPUs, and ensures on-premise…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

KNOWS Benchmark: Evaluating Web Agents on Knowledge Synthesis and Organization

Introduction of the KNOWS benchmark for evaluating web agents on complex, open-ended tasks. Agents achieve <3% success rate on full tasks; failures in visual steps prevent outputs, even when agents complete >50% of the other steps. The benchmark identifies limitations in tool use, visual understanding and…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Review of Fake Review Detection: From Traditional ML to LLM Models

The scientific review covers 211 studies (2018–2026) on fake review detection. It tracks the evolution from traditional machine learning through PLM models to LLM-based approaches. It analyzes eight types of evidence sources (text, sentiment, behavior, metadata, graphs, multimodal content). It identifies open challenges in adversarial…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

ViSTA: an adapter for clinical analyses in multimodal language models

The study presents the ViSTA adapter (0.516M parameters) for integrating clinical measurements into pretrained vision-language models. On MIMIC-IV, it achieves an AUC of 0.7376 for predicting acute kidney injury with 2B parameters (comparable to GPT-5.6 Sol) and 69.27% accuracy on temporal questions with 90% fewer…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

LUMO: An offline voice assistant with a compressed LLM for edge computing

A research team presented LUMO, an offline voice assistant running on a Raspberry Pi 5 with 8 GB RAM. The system combines local ASR, a 4-bit GGUF quantized LLM, and TTS. It achieves a WER of 6.8%, latency of 2–4 s, and power consumption of ~9 W. It supports English and Bengali without a cloud connection.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Automatic taxonomy of failure modes in retrieval-based verification for medical LLMs

A study on arXiv analyzes where LLM-generated medical answers verified against authoritative sources fail. The authors induce a taxonomy of retrieval errors (5 dimensions) and reasoning errors (6 steps) using LLM-as-Judge, tested 4 retrieval methods and 6 frontier models, and found that scaling…

arXiv cs.CL (Computation and Language / NLP) Original source ↗