A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
845
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
The EvoCause study uses LLMs to improve causal graphs for root-cause analysis (RCA) in telecommunications. It presents results on synthetic data (Node F1 higher by 11.59 pp) and the new TeleRCA benchmark with 485 681 alarms. Missing alarm name information reduces performance by 6–8 pp.
DoTime was released as an open-source PyPI package with four evaluation suites for testing causal inference in time series. The tool generates synthetic structural causal models with continuous-time interventions, counterfactual sampling, non-stationary dynamics and deterministic profiles (ramp…
The research team introduces MiGUE-Bench, a benchmark for evaluating LLM capabilities in event analysis with four tasks: detection, relations, structure and prediction. It includes MiGUE-Pipeline, an LLM-driven framework for automated annotation. Tests on existing models revealed critical shortcomings in cross-document…
The research team developed a method for analyzing intertextuality in classical Chinese histories. A large language model identifies where texts repeat and classifies them along five dimensions (form, aspect, source attribution, function, stance). On a benchmark of 2 533 pairs from the Analects and the Book of Han, models achieve…
A Microsoft research team introduced EvoLib, a framework for continuous learning by AI agents. Instead of storing raw experiences, the system extracts and iteratively refines knowledge from experience. In tests on various tasks, EvoLib outperformed memory retrieval approaches and made more efficient use of…
A new research paper describes two approaches (insertion-based and latent-space masked diffusion) intended to enable masked diffusion models to truly generate in any order, including code infilling. The authors trained a 7B model FlexMDM on Python code and a 125M model LatentMDM on GSM8K, the code is on…
Researchers introduced the DuplexGen framework, which calibrates the generation of training dialogues based on human preferences so that AI adapts turn-taking to the specific scenario. In tests on six tasks, it aligned with human preferences more closely than existing methods.
Researchers described RAGuard, a two-layer defense for RAG systems against corpus poisoning: an adversarially trained retriever and the ZKIP filter, which, according to the authors, reduced the attack success rate to 0.000 without labeled data while maintaining retrieval accuracy.
The EC-Reason-Bench benchmark shows that general-purpose language models without external supporting information all but fail at determining detailed EC numbers for enzymes, while access to external evidence sharply improves performance and blurs the differences between models.
Pegasus is a framework for transforming human manipulation videos into data for robot learning. It bridges the gap between human and robot morphology (embodiment gap) through structured knowledge transfer via Task, Affordance and Constraint Graphs with physics-based verification. Evaluation on GTEA Gaze+ and…
The study describes GuideSkill, a system that converts clinical guidelines into executable rules for LLMs. According to the authors, it improves diagnostic accuracy over RAG and model fine-tuning without changing the model weights.
Researchers introduced DuPLeR, a framework combining multimodal LLM signals with structural reasoning for knowledge graph completion in settings with limited data. This is academic research with no information provided on availability or licensing.
A study on arXiv examines lossy verification schemes for speculative decoding in LLMs and shows that accelerating generation can lead to a significant decline in output quality if the schemes are configured poorly.
Scientific research introduces AgentMap, an LLM-based multi-agent framework for ontology matching that combines the identification of equivalent concepts and subsumption. It uses semantic search and hierarchical search. It outperforms baseline methods on four OM datasets in hybrid, equivalence-only…
A research paper from Apple Machine Learning Research introduces the MoMo framework. A two-stage imitation learning model with a transformer enables robots to perform manipulation tasks with variable motion modes. Across six tasks, they achieved generalization to unseen combinations of tasks and motion modes.
Scientists developed a neural network with data augmentation to classify entanglement structures in continuous variables. The method combines classical data processing with quantum principles. Tests on tripartite and quadripartite states showed a significant improvement in accuracy.
A study on arXiv shows that while language model quantization (e.g. to 4 bits) reduces the degree of training data memorization faster than general model capability, it still allows large models to reproduce most memorized sequences verbatim – quantization therefore cannot be considered a safeguard…
The MyoCardBench benchmark, with 2 263 tasks drawn from real cardiology data, tested seven LLMs. GPT-5.4 scored highest, followed by Gemini 3.1 Pro and Qwen 3.6 27B. The models failed most often when reading ECGs and on medical ethics tasks.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.