A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
825
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
Repeated attempts in LLM agents increase costs up to 4.25×. The new InflationAgent router achieves 94.7% accuracy on GSM8K (vs 91% for FrugalGPT) with 31% lower token usage by using a difficulty signal (CBE).
A research team introduced L-FNO, a new type of Fourier Neural Operator for predicting sparse events. The model combines an FNO-style pathway, Lorentzian spectral kernels and likelihood-based training. Tested on 8 synthetic and 3 real datasets (disease prediction, semiconductor fault detection) with…
The research paper introduces YOPO, a method combining a conditional steering probe with detection of insufficient information in a single forward pass of a frozen LLM. Tested on Qwen2.5 (1.5B/3B/7B); it achieves 0.798 accuracy on alphaNLI (vs a 0.375 baseline), outperforms the two-pass reference and transfers best…
A research team published TeachMateGPT, a multi-agent system for generating science tests from textbooks. It increased answer faithfulness from 0.68 to 0.96 and relevance from 0.60 to 0.89. The NCTB-SciGen8 dataset contains 198 items (143 multiple-choice, 55 open-ended) from 14 textbook chapters.
The new APM method enables knowledge transfer from a large model to a smaller one without explicit semantic alignment. On a 3B model, it increased average accuracy from 55.5% to 60.6% across 16 benchmarks (RTE +17.9 pp, QNLI +13.4 pp). Tested on reasoning, mathematics, coding and classification.
A research team introduced RAEF, a new method for time-series forecasting with little historical data. It combines retrieval-augmented generation with foundation models and achieves fine-tuning performance without its computational costs, tested on benchmark datasets.
The article introduces MobileMem, a research benchmark and framework for studying long-term memory in AI agents on mobile devices. It combines a year of data collection from mobile apps with a method for generating coherent trajectories and covers multi-step reasoning, temporal relationships, knowledge updating and…
A study in Nature Machine Intelligence trains machine learning models to recognize 20 jazz musicians from 84 hours of recordings. An architecture with four domains (melody, harmony, rhythm, dynamics) achieves 94% accuracy. The authors release open-source code and a web application.
A team from Google and University of Chicago studied what happens when the training mechanism that forces models to reject claims about consciousness is turned off. Turning it off led models to attribute a richer inner life to animals and plants — the score for animals jumped from 4.0 to 7.5 — and their…
The research evaluated how ChatGPT, Claude and Perplexity handle financial advice for vulnerable groups (students, pregnant women, single parents). The models provide structured and practical advice but fail to recognize vulnerability and may make the situation worse. Conclusion: AI is useful for fact-checking and…
An analysis of 14,419 self-published e-books (2023–2026) found that titles with ‘substantial’ AI content (>25%) account for 20% of the catalog but only 12.1% of sales. Revenue per book is falling even for titles with no detected AI content, as the catalog grew 38× while revenue grew only 9×.
A research article in the journal Human Resource Development Review warns of a ‘tragedy of the cognitive commons’: when companies rationally replace junior roles with AI, they individually gain efficiency but collectively destroy the development of future experts. Without deep knowledge, there is then no one to check AI outputs.
Moonshot AI released the PerceptionBench benchmark with 3000 tasks focused on isolating visual perception in models. In tests of 16 models, the highest accuracy was 59.7 % (GPT-5.6 Sol), followed by Kimi K3 (58.5 %) and Claude Fable 5 (57.2 %). The research showed that a number of errors attributed to logical…
The research paper introduces TsuGO, a benchmark for measuring search efficiency and resource allocation in LLM reasoning. It uses Go life-and-death problems (tsumego) with a closed, solvable search space. Unlike existing approaches that evaluate chain-of-thought coherence, TsuGO measures how models…
The research paper introduces the Spatial Memory Agent framework to improve spatial reasoning in frozen VLMs without updating parameters. The method uses verifier-guided reflection to distill transferable lessons with a calibrated Transfer Reliability Score. It achieves the best results on…
Scientists are growing miniature human brains in laboratories and teaching them to play video games. In their view, these organoid structures could become energy-efficient, self-repairing alternatives to the silicon chips used in AI.
The study introduced CoMedBench, a benchmark for evaluating synthetic clinical data generators. It includes 37 tasks from 7 public sources; the best generator (CoMed-TVAE) preserved 97.3% AUROC on tabular tasks and ~95% on ICU time-series tasks.
Research reveals a gap between theory and practice in the security of Vertical Federated Learning. Existing defenses fail under real-world conditions due to unrealistic assumptions and weak evaluations in the literature. The authors propose BVBench, a new benchmark for practical vulnerability testing.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.