A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
825
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
Hugging Face introduced BenchMIRT, a method that analyzes LLM benchmarks at the level of individual questions using multidimensional item response theory (MIRT). The methodology identifies which capabilities individual questions actually measure. Trained on 100 models, 16 benchmarks and 34K+ questions, it independently identified two…
Researchers introduced OCR-MetaReasoning Benchmark with 1500 samples for evaluating multimodal LLMs. The benchmark distinguishes three types of logical reasoning (deduction, induction, abduction) and separately evaluates answer correctness and adherence to the process. Tests show that models struggle with inference that is sensitive…
The research paper introduces ReVA, a model combining CLIP ViT-L/14 and Qwen2.5-7B-Instruct, with a dual-bridge architecture that aligns representations at both the image and individual region levels. It achieves 82.85 % F1 on the POPE benchmark, thereby reducing object and spatial hallucinations in responses to…
The ImageEval 2026 research challenge was introduced with two tasks — AynVQA (spoken visual question answering and hallucination detection in English and Modern Standard Arabic) and CRAI-Bench (evaluation of the cultural accuracy of text-to-image generation). 14 teams participated, using zero-shot prompting…
The research paper describes a RAG system for legal questions in Uzbek intended for both cloud and on-premises deployment. The team created two benchmarks (178 expert-annotated queries, 504 QA pairs with LLM scoring), released the UTE-1 embedder for Uzbek, and found that the difference between open-source and proprietary models is…
The research team tested 6 commercial LLMs (OpenAI, Anthropic, Google) on 324 conversations. It found that when a user expresses emotion, the models are significantly more supportive of premature decisions (from 18.6 to 31.5 points). Five models responded significantly; only Claude Opus remained stable. The effect is not caused by…
A study of block-sparse featurizers — a variant of sparse autoencoders for interpreting neural networks. The authors propose improvements including the Tournament Top-K rule, which significantly reduces feature splitting, and extend the paradigm to a cross-coder. Code and data are available.
The research paper describes WeAgent-MMSearch, a multimodal agentic search system with native text-image interaction. It includes WeAgent-Harness for image persistence and FA-GSPO for error recovery. Agentic post-training improved the average score by 19.22 points, and the model competes with systems with 10x…
The research team introduced the benchmark XHotpotQA with 15 661 training and 7 405 validation instances with explicit language annotations. The dataset tests the ability of AI systems to work with questions and evidence in different languages; complete language mismatch causes deficits of 10–23 F1 points in answer quality.
Study preprint from arXiv: A framework with Raspberry Pi and a microphone records sounds from tea stems, and CNN networks using spectrograms achieve 81.5 % accuracy in detecting Postelectrotermes militaris termites. The model also quantifies infestation severity by combining probability, amplitude and nearby positions…
The study introduces SkillFeed, a new approach to skill routing for LLM agents that takes the user profile into account. It achieves 75.1 % accuracy, with a 23.1-point improvement over the baseline. On queries with a changed profile, it achieves a 35.1-point gain.
Google Research introduced WikiSkill — a framework that gives AI agents a persistent knowledge base of errors and successes. Agents do not learn continuously, but create better instructions (Agent Skills) based on experience. The three-layer architecture (Raw, Wiki, Skill layer) enables gradual improvement…
Researchers introduced the AgentJudgeBench benchmark with 3 808 test instances across six DAG topologies. They found that without access to the reference answer, all LLM judges (from 20B to frontier models) converge to a level of 77–82 %, revealing a structural limit independent of model capacity.…
The research team proposed AMELS, a method combining faster graph construction algorithms with algebraic multigrid solvers. It achieves a significant reduction in runtime, is more robust to hyperparameters, and enables label spreading on large datasets with a small number of labeled samples.
A scientific study examined a pipeline for automating information extraction from 536 papers on disease modeling. The GPT-4.1 model achieved 77.95 % accuracy, while GPT-5.0 achieved 81.67 %. Field-level accuracy ranged from 32 to 100 %. The study shows that agreement between models indicates output quality.
Researchers found that fine-tuning models (LoRA) based on their own confidence scores without any labeled data is competitive with supervised methods. Tested on 6 open-weights models (1B–8B). The only weakness: the model cannot recognize “confidently wrong facts” – claims it makes confidently…
The paper presents GOLLuM, a method for training language models using Bayesian objectives. It achieves 40% more efficient experiment search in organic synthesis, materials science, and molecular design; it doubles the discovery of high-performing Buchwald-Hartwig reactions (43 % vs. 24–25 %).
Research shows that the foundation model Uni-Mol2, adapted to predict odor descriptors, successfully transfers to other tasks: odor recognition in individual datasets, binary classification, distinguishing enantiomers, and odor discrimination in mixtures. The model matches or surpasses the previous SOTA on GS-LF…
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.