A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
825
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
The research study tests predictions of glacial lake events and landslides in the Himalayas using freely available satellite data (radar interferometry, weather) and machine learning models. Using data on 589 lake outbursts and thousands of landslides, the authors compare deep learning with gradient boosting. Antecedent weather…
Nature Machine Intelligence published a systematic review of how large language models integrate audio and speech. It covers four key areas: sound understanding, audio generation, speech and audiovisual understanding. The authors analyze how LLMs transform audio perception into deeper…
A perspective in Nature Machine Intelligence points out that knowledge in LLMs forms an interconnected system. Updating knowledge without considering its interdependencies disrupts a model's logical reasoning ability. Three directions are proposed: editing deductive closure, integrating the model's beliefs and…
The research paper proposes SAG, a new approach to retrieval-augmented generation (RAG). SAG indexes documents as event-entity pairs rather than traditional knowledge graphs. It achieves the best benchmark results, with 80.36% Recall@5 on MuSiQue (11.52 points better than competing approaches).
A research team developed BEST-KAG, a system for answering questions about building standards. It combines multimodal knowledge graphs and LLMs. The system processed 251 building standards with 171,652 nodes and 310,914 edges. In tests, it achieves improvements of up to 74% over baseline LLMs.
The scientific preprint introduces FarSky, a method for forecasting solar irradiance using a latent diffusion model over a task-aware representation. It improves accuracy by up to 11 percentage points and achieves an F1-score above 60% in ramp event detection. Tested on data from Almería, Spain.
The research paper introduces TREX, a knowledge distillation framework for foundation models that solve partial differential equations. It generates synthetic trajectories from a fine-tuned teacher and exposes the student model to long-term states. Results: student models achieve the teacher's accuracy with…
A research team introduced LoongReflect, a method for training LLM agents with improved reflection in long-term decision-making. It combines knowledge distillation from a privileged teacher with GRPO trajectory optimization. Tests on retrieval with RAG and mathematical reasoning show consistent improvements over…
The research paper addresses an LLM failure (bidirectional rationalization) in recommender system evaluation. Behavioral alignment with fine-tuning achieves a 32.19% improvement in Macro-F1 over zero-shot. The approach matches the production baseline without manual overhead and offers interpretable reasoning.
A research team fine-tuned an open language model using Group Relative Policy Optimization (GRPO) with an LLM-as-a-judge reward function to generate financial recommendations. In a causal audit with a CATE estimator, the model achieves approximately 2× higher gross-profit lift ($0.0228 versus $0.0104). It has the lowest…
The research addresses Argument Language Mismatch – a phenomenon in which models call the correct tool but generate arguments in the wrong language. Supervised training (SFT) serves as a strong baseline with performance comparable to RL; methods such as GRPO bring only marginal improvements. The authors verified that careful…
The scientific study evaluates the trustworthiness of small language models (SLMs) across fairness, robustness, privacy and ethics. Quantization preserves trustworthiness better than pruning. Compressing trustworthy models through quantization produces more reliable SLMs than training from scratch.
The research revealed a critical gap in enterprise RAG: LLMs satisfy 80% of individual constraints, but only 26.8% of responses satisfy all of them simultaneously. The new EnterpriseRAG benchmark, with 983 expert-validated samples across 6 domains, tests 13 SOTA models and simulates three critical failure modes: retrieval noise…
A research team applied higher-order numerical integration schemes (Runge-Kutta 2 and 4) to iterative decoder refinement instead of scaling model size. They achieved 22.96 BLEU-4 on PHOENIX-2014-T and 19.34 BLEU-4 on CSL-Daily, outperforming the IPSLT baseline with fewer layers.
Research from Apple and Harvard proposes an unlearning method that identifies training data with negligible influence on the model and excludes it from the removal process. The approach theoretically saves up to 50 % of computational costs. It uses influence functions for analysis on both language and image tasks.
A survey of 300 leaders showed that AI has access to only 45% of corporate data on average (just 30% at lagging companies). Companies with 70%+ access trust agents 100%, while others struggle with scaling and speed. Gartner predicted that by 2027, agents will augment half of business…
A survey of 215 radiologists published in Clinical Imaging shows disappointment with FDA-approved AI tools: only 35% see fewer missed cases (59% was expected), 9% fewer unnecessary biopsies (36% was expected) and 29% less exhaustion (56% was expected). Main barriers: costs and…
Microsoft Research introduced MindTopo, a benchmark for evaluating VLMs' ability to understand topological properties (connectivity, enclosure, knotting). Models perform better at static recognition but fail at planning tasks where they must maintain their understanding throughout actions. All models remain below…
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.