A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
825
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
The research paper introduces the UniHall dataset for categorizing hallucinations in multimodal models and the SAMF framework for self-adaptive fuzzing. Experiments reveal performance degradation in state-of-the-art models under stress testing and a trade-off between usefulness and hallucinations in RL-aligned models. Code and…
A study using the OmnilingualGAIA2 benchmark in ten languages (five writing systems) tested seven frontier AI agents. All exhibit a cross-lingual gap of 8.8–18.4 pass@3 points concentrated in tool orchestration. The drop does not shrink as model size increases; in non-Latin languages, it is caused by the loss…
A multimodal AI system for damage assessment combines a locally hosted language model with retrieval-augmented generation, thermal imaging, foundation models and wireless sensing. Knowledge graph-based retrieval outperforms the vector-based approach in cross-document reasoning; multimodal fusion…
Researchers introduced ATLAS, a framework that uses LLM agents for medication safety in patients with multiple conditions. The system structures guidelines as a medication safety graph and uses targeted questions to create patient-specific conflict graphs. The GeriMedBench dataset was also created. According to…
The research article examines banking fraud detection using machine learning. The authors tested logistic regression (AUC 0.946) and stacked generalization (AUC 0.954). The work focuses on data preprocessing and handling imbalanced datasets.
A research team proposed ZeroLock, an algorithm that enables large language models to be trained on memory-constrained devices without backpropagation. The prototype reduced memory consumption by 26.5% and increased throughput by 4.9% compared with standard approaches. Its convergence differs from backpropagation by only…
The study addresses weaker supervision in self-distillation when generated sequences deviate from the target. It proposes dual supervision: one for actually visited states and another for correct contexts. Adaptive weighting adjusts to sequence quality. Tests across multiple model sizes confirm improvements…
The research paper introduces OpenVisTool, a framework for creating training trajectories for visual tool use in multimodal agents. Key insight: supervision was to be provided only by cases where observations from tools causally contribute to the correct answer. The OpenVisTool-42K dataset contains 42 000…
A research paper proposes computational argumentation as a formal foundation for Evaluative AI – an approach that presents users with competing hypotheses and evidence for and against them instead of a single recommendation. The aim is to create an explicit and contestable system to support human decision-making.
Researchers from MIT CSAIL and Tsinghua University have developed GeoPT, a new approach to pre-training for physics modeling. It enables training 2× faster and with 60 percent less data. The goal is to improve simulations for robotics, the automotive industry, and other applications.
The FineBooks project by Hugging Face and EleutherAI tested 14 open-source OCR models on 2165+ historical pages. The best models achieved >97% character accuracy at a cost of <$2 per 1000 pages and are suitable for training LLMs, not for scientific applications. Goal: reprocess 300 000 public-domain books from…
Hugging Face released a paper on more efficient knowledge distillation. The new approach (offline top-K logits and chunked KL loss) reduces VRAM requirements from approximately 250 GB on a single GPU to a fraction of that amount, enabling distilled models to be trained on one GPU instead of hundreds.
An article in MIT Technology Review argues that AlphaFold is not an ideal model for AI in science. Its success required rare conditions: 170 thousand mapped proteins, 53 years of development and 21 billion dollars. The author therefore proposes AI agents as a better route to accelerating scientific progress.
MIT Technology Review analyzes why transformers — the key LLM architecture since 2017 — are showing limitations. Their dense attention mechanism requires up to 50 million operations for 10 000 words, while OpenAI plans to spend 50 billion dollars on computing resources this year. Transformers also cannot efficiently…
A preprint on Capek 0.5 (a vision-language model for robots, with 2B and 35B-A3B parameters) presents four capability families (spatial reasoning, temporal understanding, action guidance, state verification), trained separately with reinforcement learning and combined through weight-space merging. Tested on benchmarks and in simulation.
Researchers introduced GRASP, a method for training small language models that anonymize text without sending data to third parties. It improves the trade-off between privacy and text quality compared with the DPO baseline and runs at 1% of GPT-4o's cost.
Research on detecting misinformation through activation engineering in latent space. The method was tested on 11 models (Gemma, Llama, Qwen, 270M–12B parameters) and outperforms the baseline on the LIAR and FACTors benchmarks. It requires neither fine-tuning nor an external knowledge base; code is available.
The research paper presents EntopyMoE, a new Mixture-of-Experts architecture for byte-level LLMs. Instead of using uniform computing capacity for all byte patches, it dynamically routes experts based on entropy, allowing adaptive allocation of model capacity. It achieved the lowest bits-per-byte among…
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.