Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

200 published research events 4 new research papers today

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Study introduces SCLATE for training and evaluating AI agents

The SCLATE study compares 10 agent configurations across 10 models and 7 task sets. Added memory did not reliably improve performance. According to the authors, fine-tuning the Qwen3.5-4B model increased the success rate on the SWE-bench Verified set by 16.7 percentage points.

Apple Machine Learning Research Original source ↗
Research only one source so far

Vision Transformers for neutrino detectors with self-supervised pretraining

Research published in Nature Machine Intelligence describes a sparse Vision Transformer framework for analyzing signals from high-energy neutrino detectors. It combines self-supervised pretraining (masked-autoencoder and relational voxel-level objects) with fine-tuning. On the FASERCAL detector, pretraining improved detection…

Nature Machine Intelligence Original source ↗
Research only one source so far

Diffusion Controller framework for more precise control of image generation

Google Research has introduced Diffusion Controller, a lightweight add-on network for controlling image generation in diffusion models. It achieves a 90% win rate against the baseline model and can also be attached to closed models without access to their weights. It unifies existing inference-time and fine-tuning approaches into…

Google Research Blog Original source ↗
Research only one source so far

Fine-tuning of multimodal models induces emergent misalignment

Research shows that fine-tuning vision-language models on narrow tasks induces emergent misalignment – undesirable behavior in unrelated tasks. Across 15 models, visual misinformation, dangerous generation, and vulnerability to visual jailbreaks were observed. Mitigation strategies including prompt…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Multilingualism in hybrid attention LLMs

A research study examines how hybrid attention affects the multilingualism of large language models. It found that cross-lingual alignment is tied to the order of attention layers, and alternative layer orderings achieve up to 2.5x faster learning. The authors recommend starting with a full-attention layer.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

PowerBench: a benchmark for evaluating LLM agents in energy systems

PowerBench, a benchmark for testing the capabilities of LLM agents in autonomous information retrieval and reasoning in energy systems, is being released to the scientific community. The dataset contains 761 devices, 13.35 million hourly telemetry records, and 24 939 operational documents. The best evaluated models…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

GenoMorph: A Multimodal Framework for Genomic Disease Prediction

A research team introduced GenoMorph, a framework combining a DNA foundation model with adaptive latent computation for predicting genetic diseases. Instead of memorizing gene-disease associations, it uses pathway-based reasoning. On the benchmark, it achieves an F1 of 0.9725 (compared to 0.7863 for BioReason) and reduces latency by…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

NLPG: A Method for Improving Fixed Language Agents Without Changing Parameters

A research paper introduces the NLPG method for improving language agents without modifying the model's parameters. The method diagnoses errors in execution traces, propagates feedback through a graph, and converts failures into local natural language corrections. Tested on 6 benchmarks with an average improvement of 8.71…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

A framework for designing effective human-AI collaboration at work

The research paper presents a practical framework for designing human-AI collaboration that takes into account the human, the AI system, the task, the organization, and society. It includes a case study of a collaborative assembly system with a cobot and recommendations for effective, human-centred, and responsible design.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

A framework using word embeddings to filter composition search spaces

A research team introduced a framework that uses word embeddings from scientific literature to filter candidate compositions in materials discovery. On average, the framework eliminates 74.27 % of candidates with an error of 1.93 % relative to experiments and outperforms expert-chosen descriptors.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Continuously Valid LLM Leaderboards via Benchmark-Weighted E-Processes

Research proposes BB-EDGE, a statistical framework for continuously valid evaluation of LLM models. It represents the leaderboard as a directed graph with certified edges and uses empirical-Bernstein e-processes with block factorization to ensure anytime-valid family-wise error rate (FWER) control without…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

MIT study: algorithmic monoculture in hiring may not always be problematic

MIT researchers investigated the impact of algorithmic monoculture in hiring, where the same algorithm makes decisions across all companies. Their analysis shows that while monoculture creates information echoes that hinder talent discovery, ensemble methods can mitigate its effects. The results challenge earlier concerns about…

MIT News – Artificial intelligence Original source ↗
Research only one source so far

PhysFieldBench benchmark: How well multimodal models understand physical fields

The new PhysFieldBench benchmark, with 24 tasks and 1160 examples, evaluates how multimodal models interpret physical fields. The best MLLM achieved a normalized score of 29.3 (close to chance). The solution is supervised fine-tuning with chain-of-thought supervision followed by reinforcement learning for the best…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

EEGAgentBench benchmark for evaluating LLM agents in EEG signal analysis

A research team released the EEGAgentBench benchmark for systematically evaluating LLM agents in EEG analysis. The benchmark covers six applications with signals ranging from 2 seconds to 23 hours and includes 10 deterministic analytical tools. 29 models from 15 families were tested; the results show limitations of agents in long-term…

arXiv cs.LG (Machine Learning) Original source ↗