A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
336
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
The Omni Demand Understanding (ODU) benchmark evaluates the ability of multimodal LLMs to understand contextual requirements in audiovisual interactions. An evaluation of 14 models showed that the best (Gemini 3.1 Pro) recovers only 44.7 % of key information from visual or acoustic context; 11 of 14 models have…
The CoLearn research project is developing an agentic tutor that learns about each student's state and misconceptions. The system tracks mastery using Bayesian Knowledge Tracing with a knowledge language model and generates adaptive questions targeting the weakest topics. A/B tests show a 68–69% preference…
A research team released BI-Bench, the first benchmark for evaluating LLM capabilities in end-to-end business intelligence. Frontier models achieve <50% accuracy. BI-Agent, an agent with a tool-augmented approach (search, join, transform), increases accuracy by 40 points; post-training with SFT and RL adds up to 30 points.
A research paper from arXiv compares approaches to chip design using LLM agents. The new AHRR approach (Agent-based HLS Design with Post-HLS RTL Refinement) achieves a 2.6× geometric mean speedup over direct RTL design on an 11-task benchmark. The code and artifacts are publicly available.
A scientific study describes ArenaFlow, a framework combining tournament-based relative ranking with reflective evaluation for RL agent tasks. The system propagates advantages to the level of individual steps and skill memory, enabling more targeted optimization of agent behavior in open domains.
The study adapts the DINOv3 Vision Transformer model for automatic target recognition in underwater sonars. Using LoRA increases the AUPRC metric from 0.300 to 0.679 +/- 0.027, while training only 0.26 % of the parameters. Additional refinement techniques did not yield a statistically significant improvement.
Research into using LLMs to detect misconfigurations in Kubernetes. Researchers created a taxonomy of errors, evaluated detection tools and analyzed the severity of the issues.
A research paper proposes Continuous Delayed-Memory SGD combined with continuous-time policy gradients to model variations in quasar brightness. The new algorithm exhibits better exploration and convergence behavior than standard SGD in 2D experiments.
The share of American adults who use AI almost daily (6–7 days a week) rose from 8 % in March to 19 % in August 2026. A study by Epoch AI and Ipsos using representative samples (March: 2017 respondents, August: 1016). A change in methodology between surveys (March: AI in general, August: individual services separately)…
Microsoft and University of Illinois introduced StudentSim – a system that creates digital replicas of individual students from limited data. The replicas simulate typical mistakes and learning from explanations. The two-stage training first learns from pooled data from all students in the course, then adapts to the individual.…
The RoboHarm benchmark tested the refusal of dangerous commands (stabbing a doll, chemical mixtures, hazardous electrical work) when controlling robots. The model GPT-6 Astra refused 2 out of 100 tasks, Claude Fable 5.1 refused only stabbing a doll, and MolmoAct2 never refused. No model has a reliable safety layer for…
Researchers at Google DeepMind developed the Dream-RSI method, which allows AI agents to test new search strategies against records of past runs without having to repeat expensive computations. The agent records all attempts and results, then can simulate thousands of variants without calling the model again or…
Anthropic confirmed that it operates a biological laboratory to test AI models in real-world experiments. The company did not disclose the specific research focus, but denied a focus on drug discovery. It launched the Life Sciences Verification Program for researchers.
Businesses deploying AI lack operational capacity: decision-making on models, data management and agent governance. A survey by Collibra: 72 % of AI leaders see weak data foundations behind initiative failures. A report by EY: 60 % of organizations do not know who manages agents after deployment, 50 % have not updated governance…
SemiAnalysis publishes benchmark results for the Engram technique, which optimizes vector embeddings in LLM models for more efficient memory offloading from GPU to DRAM and SSD. In benchmark tests of DeepSeek-V4.1-Flash, NVIDIA GPUs outperform AMD MI355X. The InferenceX benchmark has support from Google Cloud…
SemiAnalysis (newsletter feed — hardware, chips, AI economics)
Original source ↗
Researchers introduce Uni-LaDiR, a framework that uses latent diffusion to map reasoning steps across different modalities into a shared latent space. The method achieved a 7.3% improvement on 11 vision-language model benchmarks and 6.1% on robotics tasks.
The study examines the benefit of having a reference solution available compared with distillation without a reference, using the AMPLE-Math dataset (5 319 mathematical problems with 6 variants). Reference-free distillation is the main source of improvement in Qwen3-1.7B; the additional benefit of a reference is modest and greater for polished solutions. SmolLM3-3B…
The research paper presents TrioRAG, a method for multimodal retrieval without building graphs. It combines a question, an anchor image and a VLM-enhanced query through late fusion; it matches the performance of graph-based systems with 1.6–2.3 times lower latency. It introduces AutoQA, a benchmark for cars with web-sourced images.
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.