A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
716
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
Researchers from arXiv introduced Graph Theory Bench with 100 000+ examples (24 graph problems in four representations) and Graph Theory Agent (GTA) to improve model performance. GTA improved Phi-4 from 53.5 % to 69.1 % on easy tasks and from 33 % to 41.5 % on hard tasks. The code is freely available.
The research examines how edits to knowledge graph models (KGE) can push correct answers out of the top results, even though the edit appears successful. On the FB15k-237 dataset with the DistMult and ComplEx models, direct editing works without harm in only 23 % of cases. Support-regularized editing achieves 36–38 % harmless…
The research team introduces Benchmark Radar – a live database and search engine for AI benchmarks and evaluations. The system catalogs 1283 source records and 12916 observations from 37 data sources, including 13 direct connectors and 24 research feeds. It provides a web dashboard, a leaderboard, a CLI for offline…
The study introduces the OmniHallu framework for detecting hallucinations in multimodal models. The OmniHallu-Bench benchmark contains 10 000 samples with annotations for six tasks (image, video, audio in both directions). The optimized verifier reduces calls to experts by 66%.
The research program at Qiushi Engine focused on BabyLM 2026 Strict-Small (10 million words) involved three phases: building a frontier model, discovering principles of efficient learning, and applying them. The result was an increase in the aggregate score from 42.02 to 42.25. The models and code are available on Hugging Face and…
AWS Machine Learning Blog introduces Agent Evaluation Metric (AEM), a methodology for measuring the quality of multi-turn AI agents. Unlike holistic scores that mask errors, AEM makes it possible to distinguish between an error in a single turn and its cascading impact on subsequent turns. The metric decomposes…
Chris Schmitz presents a study mapping “agentic flooding” – an AI-driven rise in complaints and requests in public services. Analysis of 84 cases in 11 countries: UK ombudsman +170% (2022–2024), CFPB USA +500%, similar trends in Brazil and Germany. The author sees an opportunity to reform services, not a threat –…
The research team introduces a framework with Geometric Vision Parser and Symbolic Solver modules that improves geometric reasoning in LLMs without the computational overhead of multimodal models. It achieves performance comparable to the Gemini 2.5 Pro model on a benchmark of tasks from Chinese Zhongkao exams in 2025.
A preprint from arXiv introduces a probability wave framework for modeling the behavior of interacting adaptive agents. Using data from the Chinese stock market, it empirically verifies that adaptive coupled models explain 82–94 % of observed behavior (vs. <5 % for purely independent ones), and proposes their…
MatBrain has been published, a collaborative agent with two specialized models (Mat-R1 30B for analytical reasoning, Mat-T1 14B for tool orchestration). The system is lightweight and can be deployed locally, competitive with frontier LLMs. In a catalyst design test, it generated 30 000 candidate structures and…
A research paper introduced the ERASE (Erasure via Reconstructive Adversarial Signal Editing) method for selectively forgetting designated data during inference without modifying model weights or retraining. The method uses structured, class-conditioned input perturbations and maintains performance on other…
Researchers on arXiv introduced DAREBench, a benchmark for evaluating LLM models as agents. It contains 233 tasks from 22 sources organized into a matrix by input modality and execution mode. They tested 23 commercial API models and 12 open-weight models across 7587 runs. No model won in every group…
The study compared six formats for representing reasoning in LLMs. It found a discrepancy: users prefer representations based on planning and decomposition, but simpler chain-of-thought better supports verification, trust calibration, and interpretability.
Theoretical and experimental analysis of option-critic: the learned termination rule is unnecessary, policies suffer from “policy necrosis” (getting stuck on the first action), and performance improves by reducing the probability of simultaneous option failure.
Pathway is developing the BDH architecture, which performs reasoning in latent space instead of using chain-of-thought tokens. It uses Amazon SageMaker HyperPod for training and addresses transformer limitations in systematic generalization and extended reasoning without memory loss.
Hugging Face publishes research on Boundary-Aware Self-Distillation for nuanced LLM refusals. Instead of refusing entire topics, the model is trained to refuse only the harmful parts (manipulation), while answering factual questions.
The new Korean benchmark PetQA for evaluating veterinary knowledge in LLM and LVLM models contains 10 076 text-based and 8 751 multimodal QA pairs provided by veterinarians. The study evaluates 18 models using ROUGE, BERTScore and LLM-as-judge metrics with three methods: zero-shot inference, retrieval-augmented generation and…
A study by Carnegie Mellon, MIT and Cornell (472 respondents for Trump, 1035 for Kirk) demonstrated that seven-minute conversations with the Gemini model (v1.5, v2.5) reduce belief in conspiracy theories more effectively than fact sheets. The effect carried over to new related events.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.