A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
419
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
The study finds that test-time scaling for LLMs consumes significantly more energy when candidates are generated sequentially rather than in a batch. Eight sequential calls on an A100 GPU consume 4.64–4.86x more energy and have 5.77–6.12x higher latency. Measured using Phi-3-mini and Qwen2.5-1.5B on the GSM8K dataset.
The new benchmark PosteriorBench evaluates whether generative models solving scientific inverse problems actually capture the correct posterior distribution. The benchmark includes four physics problems and five metrics (maximum mean discrepancy, Wasserstein distance). The results reveal gaps in posterior calibration…
A replication study confirms that the shape of entropy in the computational chain of a language model predicts answer correctness on mathematical benchmarks. Tested on four open-weight models on GSM8K and MATH-500, with accuracy differences of +9.6 to +27.5 percentage points. The original claim about the magnitude of the drop…
The research introduces VisKG-LM, a method that compiles knowledge graphs into images offline and caches them as visual memory. A language model then uses this memory without the need for repeated encoding. On the CommonsenseQA benchmark, it improved results by 1.2 points, on OpenBookQA by 0.8 points, and on MedQA-USMLE by…
A new benchmark for evaluating language models in biophysical research, with 500 cases and 1517 tasks across 6 biological domains. The model DeepSeek-V4-Flash achieved the highest F1 score of 0.360, followed by Qwen3.7-Max (0.316) and GPT-4o-mini (0.294). The dataset is freely available on GitHub and Hugging Face.
A study on arXiv examines how Llama-3.1 (8B, 70B) and Qwen models transmit hidden information. Causal analysis shows that internal states at certain network depths control hidden information better than the geometry of output vectors. The study also reveals multi-token confusion.
The research paper introduces the RETD algorithm, which addresses stability issues in the emphatic temporal-difference learning algorithm in off-policy settings. The paper includes theoretical convergence proofs for both harmonically decreasing and constant step sizes, along with experimental validation on several tasks.
A proposed unified evaluation framework for assessing the trustworthiness of LLMs, agentic AI and multimodal systems across eight dimensions, with mapping to international standards and EU regulation; empirical validation remains to be done.
The research team published ReaLMem, a long-term memory benchmark built from authentic personal photo archives. It evaluates models at three levels: factual recall, personality inference and predictive personalization. The team proposed ChronoProfiler with temporal weighting to resolve conflicts between preferences. Tests…
A framework combining cybernetics with interoceptive inputs enables autonomous agents to monitor their internal states and autonomously update their goals. It brings together principles from artificial intelligence, robotics, neuroscience, and cognitive science.
Research from Apple Machine Learning Research introduced the DSAS method, which selectively modulates the strength of interventions that steer the behavior of generative models. The method improves the trade-off between mitigating toxicity and preserving quality without significant computational overhead and is applicable to both LLMs and diffusion models.
Pew Research surveyed 42 151 people from 37 countries (February–May 2026). In 34 of the 37 countries, the prevailing belief is that AI will lead to job losses. Concern is highest in the USA (71 %), Australia and South Korea (76 %). 41 % of respondents feel a mix of excitement and concern, 37 % are primarily concerned, 13 % are primarily…
A study of 448 five-turn conversations (in both English and Chinese) measures how AI assistants behave under repeated verbal pressure. Gemini 3.1 Pro: 50 % hard disengagement (refusal to continue), GPT-5.6 Sol: 31 %, Claude models: 0 %. Claude remained available in all cases.
A new approach to language model safety: training on 'cunning questions' – questions with misleading premises reveal hidden risks. Results: a reduction in Attack Success Rate from 17.40 % to 15.05 % and increased robustness against jailbreak attacks.
The MIRAGE research study evaluates how multimodal agents use historical information in long conversations. It found that agents generate plausible responses even without access to the history. Open-source models are heavily dependent on context continuity and do not automatically opt for alternative retrieval.
The study compared Knowledge Graph-based augmentation (G-Retriever) with standard RAG on datasets of cultural questions from Latin America. G-Retriever reduced error by 72–78 % and was competitive with RAG; the trained projection transferred to Portuguese without special tuning.
A study on arXiv introduces the agentic-eCAL metric for evaluating the energy efficiency of multi-agent AI workflows. Experiments on GPUs from NVIDIA (A100, H100) with 16 models show that transferring text between agents accounts for 0.25 % of the total energy, while the dominant costs stem from additional processing…
The research paper addresses Compression Paradox in RAG on small GPUs (NVIDIA T4): neural compression often increases latency instead of improving it. Tri-Metric Router selects among three strategies based on CPU signals without training. It achieves a 0% OOM failure rate and an F1 improvement of 5.2 points.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.