A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
846
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
The study proposes a new framework for evaluating wildfire risk prediction systems that, instead of traditional metrics, measures whether increasing risk scores consistently correspond to greater operational demands. A comparison of the expert-based DFE index, GRU models and FARS (a hybrid system with an LLM) in French…
Researchers define a mathematical "evidential ceiling" framework for measuring what red-team benchmarks can establish about AI safety under a limited budget. They found that current tests are adequate for common harms, but millions of times insufficient for rare catastrophic incidents.
A research paper combines time series modeling (Prophet) with NLP on a 6-year corpus of Bolivian newspapers. The system achieves an AUC-ROC of 0.677 for one-day forecasting and reduces Brier Score by 10.9% compared with statistical models, capturing signals of social tension in media discourse to predict road…
MissionBench is introduced with 120 missions in 5 simulated 3D environments for evaluating 22 multimodal models. The best model achieved a success rate below 35 percent (humans: 84.4 percent), illustrating the complexity of multilevel embodied tasks. Results show that larger models have better capabilities…
The LeAct research method trains models to infer chains of thought from expert system actions instead of manual annotation or distillation. On Flop Hold'em, it achieves +60 mbb/g over the baseline and outperforms direct imitation.
A research team proposed a detection system for talking-face deepfakes that synthesize video from a photograph and audio. The method extracts physiological signals (rPPG) through RhythmFormer and trains lightweight classifiers. On Celeb-DF++, a 1D ResNet achieved an AUC of 0.806; performance varies considerably…
The research tests Group Relative Policy Optimization (GRPO) on quadcopter control using the Qwen-0.5B model. Without modifications, the model collapses to a zero action (0% success), but replacing the continuous action space with a 5-way choice of PID presets achieves 98.6–100% success. The classical PID method achieves…
The scientific study describes the phenomenon of ‘context anxiety’: reasoning models have the capabilities to solve problems but fail because of premature self-doubt and poor estimates of the tokens needed. The authors show that models can learn better strategies without this problem, without having to grow in size.
A study validates the Spanish version of a psychological dependence on large language models scale (LLM-D12-SP). A sample of 386 participants confirmed a two-factor structure: instrumental dependence (using LLMs for tasks and decisions) and relational dependence (psychological attachment). Cronbach's…
Researchers introduced TriGlue, a generative model for designing molecular glue degraders. The model combines SE(3)-equivariant estimation of protein interfaces with a flow matching network for molecule generation. The authors claim it produces chemically valid structures; the code is available on GitHub.
The research team used LLMs to extract grammatical rules and examples from grammar reference books and generated synthetic parallel corpora for fine-tuning translation models. Testing on three endangered languages (Kalamang, Tuatschin, Mandan) showed an improvement of +8.8 to +3.3 ChrF++ points. The study maps…
A research approach combining chain-of-thought with latent representation. J-CoT proposes an intermediate representation based on vocabulary-indexed coefficients instead of fully verbalizing each step. In tests, it outperformed existing latent-reasoning methods on mathematical, scientific, and coding tasks.
Researchers propose a safety approach for reinforcement learning in changing environments. The system predicts whether an agent can adapt safely in time; if not, it proactively restricts its actions. Tests in a driving simulation show fewer safety violations during environmental changes.
A study examines how evaluation design affects comparisons of expert-assigned and automatic MeSH terms for classifying medical abstracts. On the Statins topic, the difference between approaches (WSS@95%) changes from +0.096 in 5-fold evaluation to +0.021 in 10-fold evaluation. In the 10-fold design, BiomedBERT and bag-of-words produce similar…
A study evaluates Gemma 4 31B-IT on 14 NRC Reactor Operator examinations from 2015–2021. Comparing 8 fine-tuning and retrieval-augmented generation configurations showed that SFT with fixed-size chunking RAG passed 8 of 14 examinations (79.7 % accuracy). RAFT lagged behind SFT; the choice of chunking strategy…
A research paper proposes HexLogicAgent, a framework that improves logical reasoning in LLMs by organizing semantics. The key finding: failures arise more from weak semantic representations than from deductive logic. Logical hexagon theory models the complete structure of semantic oppositions, which slows…
The study introduces GLASS, a framework for personalized text generation without retraining. It combines global style representations using sparse autoencoders with local style vectors for contextual personalization. Experiments on LaMP and LongLaMP data show better results than approaches based on…
Researchers introduced the SceneActBench benchmark for testing vision-language model agents capable of acting in 3D scenes. The benchmark contains five tasks with 520 cases. Testing eleven proprietary VLM configurations produced scores of 38.6–50.2 without consistent performance.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.