Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

825 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

AI adoption in organizations does not curb R&D waste

A third of companies spend 25–40% of their R&D budget on projects with no output, and nearly half of teams estimate a loss of >1 million dollars per project. According to IEEE Spectrum, companies apply AI more to analysis and modeling than to decision-making in early ideation and feasibility, where it would have the greatest effect.

IEEE Spectrum — Artificial Intelligence Original source ↗
Research only one source so far

Explainable AI in diagnosis has different effects depending on user expertise

A study by MIT and Stanford tested explainability methods (LLM explanations, heat maps) in skin disease diagnosis. Laypeople improved their accuracy, but mainly through blind trust in AI; healthcare professionals achieved better results without additional explanations. Inexperienced users are most susceptible to…

MIT News – Artificial intelligence Original source ↗
Research only one source so far

L'Oréal accelerates cosmetics development with artificial intelligence and digital modeling

L'Oréal uses AI to accelerate the development of cosmetic products: instead of physically testing hundreds of molecules, it designs tens of thousands in computer simulations and sends only the most promising ones to the laboratory. It is also switching from animal-derived ingredients to biotechnological alternatives and uses artificial skin models instead of…

Research only one source so far

The SCHEMA framework for topological evaluation of scientific agents' hallucinations

The new SCHEMA framework automatically constructs conceptual science graphs, generates test tasks (verification, multi-hop logic, explanation, coding) and detects hallucinations with topological weighting. The research found that hallucinations concentrate at key nodes in the knowledge network and that models achieve…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Chain-of-thought reasoning in low-resource Southeast Asian languages

The study introduces OSCD, an algorithm for improving chain-of-thought reasoning in low-resource Southeast Asian languages. It combines projection of high-resource trajectories with joint-embedding semantic alignment and achieves up to a 3.2× improvement on the AIME25 and HMMT25 mathematics benchmarks…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

PCSD improves agent reinforcement learning training through persistent consistency

The research paper introduces PCSD, a method for more efficient training of language model agents on complex tasks. It combines dense teacher supervision with sparse feedback from the environment. On the ALFWorld benchmark, it achieves performance 15.6 points higher than GRPO, without dependence on a specific architecture…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Measuring artificial intelligence consciousness through classical thought experiments

The research paper formalizes measurement of AI consciousness through the Conservation-Congruent Encoding framework. It separates outward behavior from internal structure measured by operational consciousness (κ_T). It shows differences between lookup and generative systems. Relevant to AI safety analysis.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

RubricReviewer: Automated peer review guided by explicit rubrics instead of direct criticism

The research team introduced RubricReviewer, a framework for LLM-assisted peer review addressing two key limitations of existing approaches: It explicitly generates rubrics as an intermediate step and combines a training-free evidence-gathering agent (Scout) with a trained review-writing model (Aligner).…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

AdaMTP: An adaptive training paradigm for multi-token prediction

The arXiv research article introduces AdaMTP, a method for adaptive multi-token prediction training in language models. Instead of a fixed horizon length, it adapts to sequence structure through entropy-based segmentation. It detects semantic boundaries and suppresses noisy signals. Tested on Llama-3.1-8B…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

GABench: a benchmark for evaluating LLM agents in graph analysis

The research team released GABench, a benchmark comprising 10 400 tasks for evaluating the capabilities of LLM agents in graph analysis. The benchmark covers three types of graphs and four task categories (retrieval, graph theory, machine learning, open-ended questions) with 84 executable tools. Experiments show that current…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Pretraining on design data: The JONES-19 study shows the effectiveness of small datasets without ImageNet

The JONES-19 research compared CNN training on design data (images from The Grammar of Ornament, London 1857). Specialized design data does not require massive pretraining; small, high-quality curated datasets are more effective. Local training with multi-crop augmentation achieved the same performance…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Kimi K3 model architecture: Kimi Delta Attention and its origins

The article explains Kimi Delta Attention (KDA), a linear attention layer in Kimi K3. It traces the technique's evolution from linear attention through DeltaNet and Gated DeltaNet. KDA improves linear attention with a delta rule for updating the matrix, addressing unbounded growth of the hidden state.

SemiAnalysis (newsletter feed — hardware, chips, AI economics) Original source ↗
Research only one source so far

A self-sustaining AI virus uses an LLM on compromised servers to replicate itself

Researchers from University of Toronto, Vector Institute, University of Cambridge and ServiceNow created a proof of concept for a virus that runs on compromised GPUs with an open-weight LLM. The virus detects vulnerabilities and creates attacks tailored to individual targets. It operates without a vendor API; the model is from 2025 and…

Import AI (Jack Clark, Anthropic co-founder) Original source ↗