Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

846 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

New study proposes a statistically grounded method for choosing a noise schedule for training diffusion models

A study on arXiv presents a statistical framework for the optimal distribution of noise levels when training diffusion models. According to the authors, an entropy-derived schedule significantly improves training efficiency in discrete domains and is competitive with standard EDM heuristics for images.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

FlowEdit: Managing conflicting LLM responses using information theory

A research paper proposes FlowEdit, a method for addressing ill-posed problems (inconsistent conditions with no solution) by controlling internal LLM reasoning flows. The method generates alternative responses for competing hypotheses. Experiments show a 68 % improvement in exact-set-match accuracy and a 24 % improvement in…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

ExecuGraph: A verification framework for generating backend code with LLMs

A study introduces ExecuGraph, a multi-agent framework for backend code synthesis with LLMs. Six agents (Planner, Code Generator, Logical Reviewer, Evaluator, Optimizer, Explainer) collaborate in a typed workflow with isolated sandbox evaluation. On DeepSeekCoder V2 Lite, it achieved an accuracy increase of 22.5…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

LeanFlow: Formalizing mathematical texts using LLM agents

LeanFlow automates the translation of mathematical publications into the Lean formal language. Tested on papers in number theory and measure theory; Kimi2.6 completes both projects within a budget of 2000 API calls, while GPT5.5 achieves the lowest token costs. It solves all five tasks from ICML 2026 AI for Math TCS…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Benchmarking personalization capabilities in large language models

A research study adapts the Bayesian Persuasion framework for generative agents and tests LLM personalization capabilities in sales. The authors released SDR-Bench, a dataset with 6 279 cases from 22 industries, and found a consistent personalization plateau — among Fortune 100 tech companies, no model differs statistically in…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Training-free routing: Managing LLM offloading through a reliability gate

A research team proposes CARGO, a training-free framework for dynamic routing between a local and a cloud LLM. The method uses prompt-varied sampling and Bayesian early stopping to estimate the local model's reliability without additional router training. In tests, it outperformed some supervised…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

VeriSimpl: Natural language optimization with simplification-based verification

A research paper describes VeriSimpl, a framework for translating natural language into optimization models using LLMs. Its innovation is "simplification-based verification" — an optimization solver generates simplified queries to verify that the formulation is correct. The method achieves higher accuracy than existing…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

From errors to rules: Iterative prompt optimization for text classification

A research study compares prompt optimization approaches (demonstration selection, exploration, error diagnosis) for text classification. It proposes ERGO, a method that iteratively diagnoses errors and generates decision rules. It achieves the best results on tasks with learnable boundaries: TREC 90…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Aligning language model expectations with real user needs

A research study examines how well large language models align with what users actually want. The authors introduce ExpectBench, a benchmark with real user expectations, and the LENS framework for generating better-aligned responses.

arXiv cs.AI (Artificial Intelligence) Original source ↗