A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
846
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
A study on arXiv shows that LLMs that perform well on single-turn tasks suffer a substantial drop in performance when users gradually reveal, refine, or change their intent during a conversation – across multiple model families.
Researchers described MiniCache, a framework that uses a small model to store Program-of-Thought programs in a reusable cache. According to experiments on WebShop, Formula and CodeTAT-QA data, it achieves up to 3.1× lower latency and 2.8× higher throughput.
The new SenCos-GEM method for molecular property prediction combines physical constraints (the law of cosines) with adaptive feature modulation and, according to the authors, achieves the best results on tasks sensitive to 3D structure in the MoleculeNet benchmark (FreeSolv, Lipophilicity, QM9).
A paper on arXiv describes a modification to the HyperGraphRAG system, which uses hypergraphs to store knowledge. The authors address errors in fact extraction using self-consistency prompting and improve passage retrieval with the Personalized PageRank algorithm applied to the hypergraph.
A new paper on arXiv describes improved SOAP and Muon optimizers for LLM pretraining, which, according to the authors, deliver more stable and higher-quality training than AdamW at large batch sizes. The code has been released as open source (NVIDIA-NeMo/Emerging-Optimizers).
The new benchmark DFAH-Bench examines whether AI agents arrive at the same financial decisions using the same process. Over 8 000 runs across 10 models showed that agreement in outcomes (95 %) significantly exceeds agreement in the process (77 %).
In a paper on arXiv, researchers described the ClickGuard browser extension, which combines transformer embeddings, linguistic features, and an XGBoost model to detect clickbait (F1-score of 91 %) and also offers a brief summary of the article.
A study on arXiv presents a statistical framework for the optimal distribution of noise levels when training diffusion models. According to the authors, an entropy-derived schedule significantly improves training efficiency in discrete domains and is competitive with standard EDM heuristics for images.
A study on arXiv describes CvAdamW—a variant of the AdamW optimizer that uses a thermodynamic analogy to monitor attention in Transformers and dynamically adjusts weight decay to accelerate grokking. On one tested task, it reduced the time to generalization by approximately 6 %.
A research paper proposes FlowEdit, a method for addressing ill-posed problems (inconsistent conditions with no solution) by controlling internal LLM reasoning flows. The method generates alternative responses for competing hypotheses. Experiments show a 68 % improvement in exact-set-match accuracy and a 24 % improvement in…
A study introduces ExecuGraph, a multi-agent framework for backend code synthesis with LLMs. Six agents (Planner, Code Generator, Logical Reviewer, Evaluator, Optimizer, Explainer) collaborate in a typed workflow with isolated sandbox evaluation. On DeepSeekCoder V2 Lite, it achieved an accuracy increase of 22.5…
LeanFlow automates the translation of mathematical publications into the Lean formal language. Tested on papers in number theory and measure theory; Kimi2.6 completes both projects within a budget of 2000 API calls, while GPT5.5 achieves the lowest token costs. It solves all five tasks from ICML 2026 AI for Math TCS…
A research study adapts the Bayesian Persuasion framework for generative agents and tests LLM personalization capabilities in sales. The authors released SDR-Bench, a dataset with 6 279 cases from 22 industries, and found a consistent personalization plateau — among Fortune 100 tech companies, no model differs statistically in…
A scientific study formalizes incomplete prompts as an attack on language models. Research shows that LLMs delay refusal until a sentence ends and that training against this type of attack fails to generalize. The authors identify neurons crucial to text completion and propose intervention…
A research team proposes CARGO, a training-free framework for dynamic routing between a local and a cloud LLM. The method uses prompt-varied sampling and Bayesian early stopping to estimate the local model's reliability without additional router training. In tests, it outperformed some supervised…
A research paper describes VeriSimpl, a framework for translating natural language into optimization models using LLMs. Its innovation is "simplification-based verification" — an optimization solver generates simplified queries to verify that the formulation is correct. The method achieves higher accuracy than existing…
A research study compares prompt optimization approaches (demonstration selection, exploration, error diagnosis) for text classification. It proposes ERGO, a method that iteratively diagnoses errors and generates decision rules. It achieves the best results on tasks with learnable boundaries: TREC 90…
A research study examines how well large language models align with what users actually want. The authors introduce ExpectBench, a benchmark with real user expectations, and the LENS framework for generating better-aligned responses.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.