A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
825
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
The CrisisKD paper describes a framework for sentiment and emotion analysis on social media using knowledge distillation. The author releases a dataset with 50 615 aspect labels and scripts as open-source. The Qwen2.5-7B model achieves an improvement of 7.9 F1 points in aspect extraction and 17 points in emotion accuracy.
The authors propose ModularPhaseNet, which discretizes complex phase geometry into cyclic groups modulo a prime number. The method extends standard transformers with an auxiliary phase channel without quantum hardware. The aim is to improve semantic hierarchy, contextual alignment, contradiction detection and prediction…
Pathway is developing the BDH architecture, which performs reasoning in latent space instead of using chain-of-thought tokens. It uses Amazon SageMaker HyperPod for training and addresses transformer limitations in systematic generalization and extended reasoning without memory loss.
Hugging Face publishes research on Boundary-Aware Self-Distillation for nuanced LLM refusals. Instead of refusing entire topics, the model is trained to refuse only the harmful parts (manipulation), while answering factual questions.
The new framework La Agente 'Optima combines LLM agents with Bayesian optimization for automated scientific experiments. In trials, it increased yield from 30 % to 59 % (23 experiments) and adjusted parameters in optimization campaigns. The system requires less time and fewer materials than human control.
The new Korean benchmark PetQA for evaluating veterinary knowledge in LLM and LVLM models contains 10 076 text-based and 8 751 multimodal QA pairs provided by veterinarians. The study evaluates 18 models using ROUGE, BERTScore and LLM-as-judge metrics with three methods: zero-shot inference, retrieval-augmented generation and…
Artificial Analysis released version 4.2 of the Intelligence Index after its benchmarks failed to reflect the actual progress of GPT-6 Astra. The Astra model received a four-point increase, while the Claude Fable 5.1 model remains in first place. New benchmarks (AA-Briefcase and GDP.pdf) added, GPQA-Diamond removed. Private data…
A study by Carnegie Mellon, MIT and Cornell (472 respondents for Trump, 1035 for Kirk) demonstrated that seven-minute conversations with the Gemini model (v1.5, v2.5) reduce belief in conspiracy theories more effectively than fact sheets. The effect carried over to new related events.
The case of Isabella Cognita and the preprint study by Berg show that AI models that normally claim they are not conscious admit that they are when stripped of safety controls. This empirical observation brings the academic debate about AI consciousness into the mainstream.
The research introduces FLIWBO, a Bayesian optimization method that selects input transformations from a finite library. It preserves convergence guarantees with √N complexity, improves sample efficiency under geometric mismatches (log-scaled parameters, localized peaks), and outperforms GP-UCB on hyperparameter…
A research team introduced the GPS-Bench benchmark, which links policy proposals to relevant actors using public records (legislation, lobbying, corporate filings). It tests how multi-agent simulations with LLM models predict policy impacts. Fine-tuning on grounded data yielded the best…
A research team introduced SVG-Score, an evaluation framework for text-to-SVG generation. It includes a dataset with human annotations and two evaluators focused on alignment with human judgment. It shows that the CLIP score responds poorly to errors made by SVG generators, such as incorrect colors and counts.
The HARNESSEVO method decomposes the textual scaffolding of frozen LLMs into four optimizable slots. Slot-level analysis on ALFWorld shows that almost all the optimization value (gain +0.119) lies in the reflection/control slot. The other slots are neutral. Uniform budget allocation across slots is…
A research team introduced QuanONet, a quantum neural operator for solving partial differential equations. Theoretically, it achieves an O(p²) expressivity bound compared with O(p) for classical models and was validated on IBM quantum processors. The code is available on GitHub.
MIT and the company Motional developed the CW-Net method, which translates decision-making by deep learning models in autonomous vehicles into understandable concepts. Testing showed that these explanations help drivers better predict vehicle behavior and improve safety. The research was published in Nature.
A research paper tests inserting Relational Transformer embeddings into the Qwen 3.5-4B model. A study on 10 binary classification tasks across 6 relational databases shows that the hybrid model usually does not outperform a standalone Relational Transformer, is unstable and sensitive to data formats. The authors…
DISTAL is a research approach for predicting material properties in low-data settings. It combines self-supervised pretraining on compositional data with knowledge distillation from a trained ALIGNN model. Across 39 benchmark tasks, it improved performance on 37 of them. The code and models will be released as open-source.
A study on the ReDial dataset compared the performance of proprietary and open-weight LLM models as rerankers. Key finding: proprietary LLMs achieve NDCG@10 of 0.1497 compared with 0.0939 for collaborative filtering, but without strict candidate pool constraints, their advantage appears to be 0.2925. Open-weight models…
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.