Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

293 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

OMatG-flash model for more efficient generation of inorganic materials

Research on arXiv introduces OMatG-flash, a flow model for predicting and generating crystalline structures. According to the authors, it achieves performance comparable to the best models, but requires orders of magnitude fewer computational steps and less machine time. It applies Reinforce Adjoint Matching for optimization.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

The SynCraft framework optimizes molecular synthesizability using large language models

The research team published the SynCraft framework, which uses large language models to predict sequences of atom-level edits to molecules to improve their synthesizability. The system outperforms baseline methods in generating synthesizable analogues with high structural fidelity and works with both proprietary and…

Nature Machine Intelligence Original source ↗
Research only one source so far

Apple proposes the Probe Guidance method for diffusion language models

A research paper presents Probe Guidance, a new method for steering diffusion language models. The method uses the frozen internal states of an existing model and eliminates the need for an additional forward pass during inference. On a model with 1.7B parameters, it achieves improved performance on multiple-choice benchmarks…

Apple Machine Learning Research Original source ↗
Research only one source so far

Antropic is building a laboratory where Claude controls robots in experiments with drugs

Antropic is building its own biological laboratory in San Francisco, where the Claude model will control robots conducting experiments with drugs with minimal human involvement. The company has developed Claude Science (an AI workspace for researchers) and Model Hardware Standard (control of laboratory equipment). In April…

The Decoder (daily AI news) Original source ↗
Research only one source so far

BabelArena: A large-scale multilingual benchmark for agents with large language models

The research team introduced BabelArena, a benchmark with 16 146 tasks across 23 languages and 702 canonical prompts. Tests of five frontier models revealed that no model dominates universally and agents in low-resource languages have higher rates of tool and control-flow errors, while consuming up to 2× more…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Personalized sleep guidance from wearable device data using language models

The research paper presents a two-stage framework for personalized sleep guidance. In the first stage, a multi-agent language model pipeline generates structured guidance from wearable device records without manual annotations. In the second stage, the approach is distilled into small models suitable for local deployment.…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

StepKV: Step-aware KV cache compression for LLM agents

The research paper introduces StepKV, a method for compressing the KV cache in LLM agents. It addresses the problem of "Reasoning Continuity Disruption" by pruning tokens at the level of reasoning steps rather than individual tokens. It combines step utility with token-level saliency and improves the efficiency-accuracy tradeoff…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

AI study: happiness sentiment rises in climate campaigns on X, but calls to action decline

The research analyzed 364 thousand posts on X from four climate actions using an AI lexical model; it detected a paradox: during campaigns, happiness rises (by 9 percentage points above baseline), but action language declines (by 10.75 points). Happier posts had lower reach through retweets.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

FLARE: Dense supervision for training AI coding agents

The research team introduced FLARE, a method for training AI agents to solve software engineering tasks. It uses a generative reward model to provide a dense supervision signal during training. According to the results, it achieves 5× lower token consumption compared with a competing approach, with a 19% improvement…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

RAG pipeline with BM25, BGE-M3 and LapaLLM 12B for offline processing of Ukrainian documents

The system for UNLP 2026 Shared Task extracts information from Ukrainian PDFs by combining BM25, BGE-M3 and Cross-Encoder reranking with the 4-bit quantized model LapaLLM 12B on NVIDIA T4 GPUs. OCR completed within the 9-hour Kaggle limit in an offline environment, with the remaining 2 hours used for LLM inference on 500 questions.…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

LE4Mob: inductive location embedding for modeling human mobility

A research team introduces LE4Mob, a framework for learning geolocation representations that preserves geographic distances and enables prediction for previously unseen locations. It combines contrastive language-location pre-training with distance-aware regularization, outperforming baseline methods in inductive settings…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

How conversational AI chatbots undermine human empathy

Academic Sherry Turkle is releasing the book "Artificial Intimacy: Who We Come When We Talk to Machines" on 29 September, in which she warns that ChatGPT, Replika and other conversational AI agents lead people to misinterpret machine behavior as care, resulting in a decline in their capacity for empathy.

404 Media (investigative tech journalism) Original source ↗
Research only one source so far

Researchers propose a framework for testing medical AI modeled on drug approval

Scientists from University of Bristol propose a "Learning Ensemble" framework for testing medical AI in three areas: system limitations and data, reliability across patient groups, clinical suitability. Inspired by the pharmaceutical process. It addresses problems where AI learns from irrelevant patterns and fails for…

The Decoder (daily AI news) Original source ↗
Research only one source so far

Pruning LLMs using Ising optimization

Hugging Face publishes a research paper that formulates block removal from LLMs as an Ising glass optimization with all pairwise interactions. At 50% compression of the Llama-3.3-70B-Instruct model, it achieves a score on MMLU that is 23 percentage points higher than competing block removal methods.

Hugging Face Blog Original source ↗
Research only one source so far

Competition among open models: Chinese companies take the lead

Author Nathan Lambert analyzes competition among open models. According to him, Chinese companies (Alibaba with Qwen, DeepSeek, Kimi K3, GLM-5.2) have led since July 2025 and have surpassed American Llama. He distinguishes open-weight from true open-source. GLM-5.2 and Kimi K3 have achieved agentic capabilities comparable to Claude…

Interconnects (Nathan Lambert, AI2) Original source ↗
Research only one source so far

SpecOpt: An Agentic Framework for Optimizing Molecular Binding Specificity

The research team introduces SpecOpt, a method for optimizing existing drugs using an agentic framework with an LLM and protein docking. On a dataset of 915 compounds from ChEMBL, the method improved binding specificity in 84.8 % of drugs, shifted the average binding gap from −0.72 to +0.47 kcal/mol, and maintained similarity to the original…

arXiv cs.AI (Artificial Intelligence) Original source ↗