A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
846
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
A scientific study tests interventions to increase diversity in LLM opinions. It found that simply increasing the temperature and providing instructions are ineffective; structured persona conditioning combined with interaction architectures is effective. The results are based on a factorial experiment with 100 questions and 7 models.
Research shows that LLMs are biased by the text in the corpus they were trained on. The authors propose a methodology for measuring bias using prediction markets as an external reference. In an experiment on 111 Ukrainian markets (93 000 predictions, four models), they found that English-language news coverage systematically introduces bias…
An arXiv study examines how 19 large language models frame political and cultural entities differently in the post-Soviet region. Using Valence-Arousal-Dominance profiling, it identified three model clusters with different affective behavior, with the rate of generic responses strongly correlated with lower…
The THOR research framework addresses two key problems in multi-hop question answering: attention decay during long reasoning chains and error accumulation. Inspired by theta-gamma oscillations in the brain, it separates global planning from local retrieval and introduces a verification mechanism to stop…
Research tests the ability of GPT-4.1 to simulate personas and predict opinions. It correctly predicted election results in 8 of 9 US states and achieved 0.94 accuracy in predicting views on childhood vaccination. Persona simulation showed promise, but lacks natural fluency.
Research results on TextGrad applied to agents. Handwritten policies improved the success of frozen 7B agents by 5.0 points, but policies generated from data performed no better than fixed prompting. The main challenge is reliably generating and selecting policies from experience, rather than executing them.
Research maps how four open-weight LLMs process animacy (the distinction between living and non-living entities). Through circuit discovery and ablations, researchers identified causal mechanisms responsible for this behavior, but found that they are not strongly localized and generalize only partially across models.
A research paper describes an LLM framework with a human-in-the-loop approach for detecting cutaneous adverse effects in clinical records. It achieved better accuracy (F1=0.88 vs 0.77), greater inter-rater agreement (Cohen's kappa 0.82 vs 0.50), and halved the time required.
LLM-INSTRUCT won the ArgMining 2026 competition for analyzing arguments in UN and UNESCO documents. It combines metadata-aware dense retrieval, constraint decoding with multi-agent debate, and JSON schema validation. It ranked 1st in F1 score; the paper and code are publicly available.
A research paper introduces AlphaAgent, a framework for analyzing literature in materials science. It separates retrieval from report generation through explicit skill contracts. In a blind evaluation on 40 materials science questions, it substantially outperformed the baseline, especially in mechanistic…
A study examines five diversity metrics in LLM ensembles across approximately 32 thousand subsets of 30 models. It finds that diversity measures model capability rather than diversity itself, and majority voting outperforms the best member in only ~10 % of cases. The results suggest limited usefulness of existing…
A study examined how DeepSeek and Qwen QwQ models evaluate literary quality. On a benchmark of 30 texts in 6 categories, they achieved 79.3 % accuracy. The models prefer intentionality, craftsmanship, depth and a distinctive voice. A second study showed that degrading structure and voice has a greater impact than simplifying lexical…
An arXiv research paper argues that natural language should not completely replace formal languages for specification. The authors introduce task specificity and prove a specificity crossover theorem — there is a threshold beyond which expression in natural language is more demanding than direct formal…
A theoretical preprint paper formalizes the concept of Semantic Field Theory (SFT) for lexical semantics. It defines semantic fields as mathematical structures with deformations, interaction complexes, and Gaussian representations. It introduces energy minimization for stable interpretation. It is purely…
An arXiv research paper presents AsymVerify, a methodology for classifying responses as clear, ambiguous, or refusals. The system achieved a Macro F1 score of 0.85 in SemEval-2026 Task 6 (second place among 41 teams), with asymmetric verification reducing errors at the boundary of ambiguous responses with…
Scientists published a Nature study in which AlphaFold AI identified key regions in gene-editing proteins responsible for unwanted edits. Modifying these regions reduced safety problems in gene therapy.
The new ATM mechanism enables multi-agent LLM systems to split overloaded agents into subagents at runtime with safety guarantees. According to the authors, it increased the success rate on programming tasks from 3.3 % to 61.7 % and reduced sensitive data leakage to zero.
A scientific paper introduces Telco-GAIA, a benchmark for evaluating AI agents working with real telecommunications data. It includes 100 bilingual questions (English, Arabic) requiring multi-hop reasoning over HTML, an SQL database and web archives. The best models solve 71 % of tasks; visual…
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.