Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

825 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Why LLMs are stagnating in long-form nonfiction writing

An author from AI2 claims that LLMs have improved at short-form writing (copy), but are stagnating in long-form, factual writing. Organizing knowledge is the compression required for insight; however, models increase entropy in long-form writing. This signals that LLMs are not ready for open-ended scientific problems, although the author remains…

Interconnects (Nathan Lambert, AI2) Original source ↗
Research only one source so far

AI-designed bacteriophages raise biosecurity questions

Researchers used the Evo 2 model to design genetic sequences for bacteriophages, some of which proved functional in the laboratory. The study raises the question of whether biotechnology safeguards are evolving quickly enough as AI capabilities grow. The model was developed with safety restrictions…

The Conversation — Artificial Intelligence Original source ↗
Research only one source so far

Recall failures, not empty shelves: Google maps the factuality bottleneck in frontier LLMs

Google Research introduces a knowledge profiling framework that distinguishes encoding failures from recall failures. An analysis of frontier LLMs on the WikiProfile benchmark (2150 facts) shows that models encode almost all information but struggle to retrieve it. Chain-of-thought and thinking mode provide a solution.

Google Research Blog Original source ↗
Research only one source so far

Measuring LLM credibility using Cross-Contextual Consistency

A research team introduced Cross-Contextual Consistency (C3), a method for measuring the credibility of large language models. It compares model responses to the same task with different, semantically neutral contextual variations. A study of 26 models and six benchmarks (reasoning, factuality…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

LLM Agents Factory: a framework for retrieving domain-specific agents

The research paper introduces LLM Agents Factory, a framework that constructs agents from a collection of over 20K profiles through semantic retrieval or distillation into a more compact model. On the MMLU and BIG-bench benchmarks, it achieves higher accuracy than the non-agent baseline and lower costs than AutoGen.…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Comparing fine-tuning strategies for locally deployable small language models in clinical decision support

A research team benchmarked 8 open-source SLMs with four fine-tuning strategies (zero-shot, prefix tuning, LoRA, full fine-tuning) on 2 083 MIMIC-IV-ED cases. Tasks: triage prediction, specialist referral recommendations and diagnosis. LoRA fine-tuned SLMs outperformed Claude Haiku 4.5 and Claude Sonnet 4.5…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Robustness of large language models to Vietnamese dialects

The new benchmark VialectBench measures how well LLMs handle Vietnamese dialects. The research team tested 10 models on 2400 dialectal variants. The average performance drop was 2.82%, with some dialects (PNT3, PNT2) causing degradation of up to 6.17%. The answering task saw the largest drop in performance…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Mapping the evolution of large language model behavior

The research analyzed the behavior of 32 language models from 6 families using responses to 10 000 prompts and three distance metrics. It found that models cluster by family, distances between families decrease over time and newer reasoning-oriented models have more compact response clouds.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

SpaHybGen framework for universal robotic object grasping

A research team published SpaHybGen, a framework combining neural networks with analytical planners for robotic grasping. The system generalized to 7 different robotic hands without retraining, with a success rate of 94.3–98.0 %. The code and models are open-source.

Nature Machine Intelligence Original source ↗
Research only one source so far

CARE-X: a Microsoft Research model for radiological analysis

Microsoft Research introduced CARE-X, a research model for radiological analysis that combines generative and discriminative capabilities to create reports, assess the presence of findings and locate them on chest X-rays. The model is not approved for clinical use.

Microsoft Research Blog Original source ↗
Research only one source so far

Stanford designed an AI-created bacteriophage that functions in the laboratory

Researchers at Stanford University used AI to design the DNA of bacteriophage ΦX174. The synthetic DNA was assembled in the laboratory and produced a functioning virus. The study in Science demonstrates that AI can learn biological systems well enough to design functional organisms. The aim is to use a similar…

The Conversation — Artificial Intelligence Original source ↗
Research only one source so far

Agentic LLM for colorectal cancer treatment planning reached the level of oncology experts

A team at the University of Florida healthcare facility developed GatorOnco, an agentic LLM trained on 282 billion tokens of biomedical text. In a blinded evaluation by five oncologists, the model outperformed open-source LLMs and achieved performance comparable to human experts in accuracy and safety, outperforming them in…

arXiv cs.CL (Computation and Language / NLP) Original source ↗