A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
846
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
The MyoCardBench benchmark, with 2 263 tasks drawn from real cardiology data, tested seven LLMs. GPT-5.4 scored highest, followed by Gemini 3.1 Pro and Qwen 3.6 27B. The models failed most often when reading ECGs and on medical ethics tasks.
A study on arXiv introduces the Right-sizing Recommendations (RSR) framework, which uses bootstrap conformal prediction to estimate future virtual machine usage and is intended to improve the accuracy of sizing recommendations at hyperscalers.
The model Claude Mythos Preview from Anthropic spent 60 hours searching for new cryptographic attacks at an estimated 100 000 dollars in API costs; the key human intervention was encouraging the model not to give up.
Simon Willison — AI tag (leading independent LLM commentator)
Original source ↗
A Google Research study analyzing 15 million work-related interactions with the Gemini model found no evidence of mass automation of white-collar work – according to the company, AI use remains shallow and predominantly collaborative.
The arXiv study introduces an efficient method for auditing LLM training data through data influence scoring without retraining. The authors applied the approach to HelpSteer2 and Anthropic HH-RLHF datasets, identifying types of labeling errors and safety problems in benchmarks.
An arXiv preprint shows that dynamic benchmarks for multimodal fact-checking are less uncontaminated than assumed. Approximately 17–29% of claims published after models' cutoff dates contain previously available knowledge, increasing system performance by up to 11 points and distorting evaluation.
The scientific paper proposes QFedPolyp, a federated learning framework with 8-bit quantization for polyp segmentation. It achieves a 4-fold reduction in communication between servers while maintaining accuracy and increasing inference speed. Tested on the Kvasir-SEG and CVC-ClinicVideoDB datasets.
A new scientific benchmark, StanceFlip, for analyzing dynamic stance changes in multimodal dialogues. It includes the extraction of a six-component stance (subject, target, emotion, sentiment, stance, reason) and the attribution of changes. The authors propose the ConStaFF framework with a Thought-of-Stance methodology that achieves state-of-the-art…
The research team combined transcription-based LLMs with speech analysis on Japanese speed-dating conversations to predict attraction. Combining both approaches improved pairwise ranking with statistical significance, while gains in Pearson's r varied by situation and participant.
Apple published a technical study on the audio synthesis architecture in Siri Expressive Voices. The system uses Diffusion Transformers with separate temporal and depth components on the AMX chip, just ~21MB of memory and execution 16× faster than real time. The MOS score improved by +0.28 (4.15 vs 3.87).
METR introduced the metric “expenditure horizon” to compare the costs of AI agents and humans on the same task. In the NanoGPT speedrun test, older models have so far delivered only a fraction of the value of human work; newer models such as Opus 5 are absent from the test.
A new study introduces the FaithC4 benchmark (1 455 documents, English, Chinese, Korean) and shows that general-purpose vision-language models performing OCR on damaged text often replace an unreadable word with a more probable alternative instead of transcribing it faithfully, while specialized OCR models are significantly…
Researchers introduced UOWReg, a regularization technique for self-supervised learning and JEPA models intended to suppress the encoding of biases based on sensitive attributes. According to the authors, it reduces violations of the Equalized Odds metric on the CelebA benchmark while maintaining comparable classification accuracy.
A preprint on arXiv claims that LLMs trained using reinforcement learning perform better when merged (model merging) than models trained using supervised fine-tuning — they have fewer task conflicts and a smaller drop in performance.
The authors of an arXiv preprint prove that a sufficiently accurate neural network approximation of the score function implies that the distribution generated by the reverse diffusion process is close to the target data distribution (in KL divergence), and derive an explicit error bound.
Researchers introduced HierFlow, a method for automatically designing workflows for multi-agent systems with large language models without training. According to the authors, it outperforms comparison approaches on question-answering, mathematics, and code generation tasks while maintaining efficiency.
A new study on arXiv proposes evaluating LLMs based on how other models express preferences for their responses in anonymous voting, instead of comparing them with a correct answer. However, the authors caution that this reflects agreement in preferences among models, not objective correctness or agreement with human judgment.
Researchers described a sparse attention method that uses the gzip compression ratio instead of trained parameters to select important text blocks. On PG-19 with 92M parameters and an 8K context, it reportedly outperformed dense attention as well as BigBird and Longformer.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.