A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.
825
published research events
Latest work
A significant claim from a single source is published only after further confirmation.
A research study uses telemetry data from connected vehicles in Sydney to proactively detect and predict high-risk driving locations. Eight predictive models (Random Forests, XGBoost, LightGBM, LSTM, N-BEATS, ARIMA, Exponential Smoothing, Prophet) were compared; ARIMA achieved the lowest error (MAE…
The scientific paper presents OraclePhys, a framework for systematically fine-tuning models for structural mechanics. A benchmark with a finite-element oracle objectively scores responses. A study across 7 response formats found that the format (not the volume) of data determines learning: a ranking objective creates a forward model, scalar…
Researchers released BEAR-Bench, a benchmark with 1000 questions in English and Russian on text-heavy business and scientific documents. Tested on 16 models, including Gemini 3.1 Pro, Qwen3.5-397B and others; hallucination detection methods were also evaluated.
The TokEval evaluation suite measures linguistic and structural properties of tokenizers (UTF-8, digit alignment). Controlled model pretraining experiments showed that information-theoretic metrics predict language performance (Spearman up to 0.80), while structural metrics correlate with tasks (linguistics…
The authors tested 5 benchmark suites on 26 open-source SLMs and found that approximately 60 % of responses are ambiguous or irrelevant. Model rankings change significantly depending on how ambiguity is handled.
The preprint presents GxP-Agent, an agent system with a DAG topology for generating CDISC datasets in clinical trials. The model Claude Sonnet 4.6 achieved 100 % structural compliance on CDISC-Bench (49 variables, 254 records), while all single-agent and flat multi-agent approaches failed (0 %).…
A research team introduces SGHA, an automated system for discovering research problems running on a local 9B language model. The system structures scientific literature into evidence-linked objects, detects unresolved patterns, and generates research problems with supporting rationale. A comparison with AI Scientist-v2 shows that…
The research shows that modality preferences in vision-language models change with the task: text dominates in arithmetic problems, while images dominate in chart analysis. The effect is confirmed in models evaluated on GSM8K, SVAMP and ChartQA, including API models.
An analysis of the Voynich manuscript tests three traditional assumptions: glyphs as letters, tokens as words and spaces as separators. Results with statistical controls show that all are wrong — glyphs have higher regularity (2.7 bits of entropy), tokens are weakly predictive (<1%) and spaces are…
A research team proposed MD-SigLIP, a method for directly aligning brain signals with text representations in a shared semantic space. It enables retrieval-based decoding instead of language model reconstruction and achieved the best results in tests on the full vocabulary and restricted subsets.
Scientists introduced VITAL, a deep learning method for predicting peptide-protein interactions. It combines protein language models with a geometric encoder, achieves an AUC of 0.87 and maps interfaces with >60% accuracy. Available as open source on GitHub and as an interactive web server.
An Apple Machine Learning Research study analyzed 21 000 conversations from four models (GPT-4o, GPT-4.1-mini, Claude Sonnet 4.6, Gemini 2.5-flash). It found that human-like behavior (self-reference, relationships, boundaries) is widespread in LLMs and varies across models and user factors. Users…
Artificial Analysis released the Search Index benchmark, measuring the performance of search APIs (Parallel, Exa, Firecrawl, You.com, Tavily, Keenable, Brave) for AI agents. Parallel, Firecrawl and Parallel turbo offered the best value for money; better quality reduces an agent's overall costs.
ALTK-Evolve research shows that agents can learn from their own trajectories and distill guidelines for future tasks. The right amount of memory varies by model: strong models benefit from the full collection (DeepSeek-V3.2 +9.5 pp), while weaker models benefit from a compact core with retrieval (gpt-oss-120b…
The MIT CSAIL research team identified the phenomenon of "attribution decay" - in models trained on large datasets, individual training data cannot be linked to specific outputs. The study shows that removing one or more examples from the training data does not change the generated content. The results are published…
Canadian researchers developed GenAI-dentity Survey and Workshop, which examines educators' readiness for generative AI. Over 800 teachers completed the survey; the research shows that they feel unprepared and unsupported, which prevents them from mentoring students in the ethical use of AI.
A study in PNAS examined X's algorithm among 715 US users (September–October 2024). The algorithm amplifies content that conflicts with personal values and asymmetrically serves polarizing content according to political orientation.
Stanford's AI Observatory project analyzed 7 datasets of real conversations with AI models. It found that companies filter data: Anthropic would filter out 48% of conversations; filtered-out conversations contain more health and relationship topics (44.2% vs 31.2%), sexual content (16.7% vs 2.4%) and hate…
AI Radar monitors Czech and international sources every day, looking for changes that truly deserve attention.
MonitorsOfficial AI company blogs, specialist media, and research sources.
Selects and combinesFilters out information noise and combines articles about the same change into a single event.
Summarizes and explainsExplains significant events in English: what happened, why it matters and where the information comes from.
The result is a quick overview of what has actually changed in the AI world, rather than another stream of articles.
Use the CS/EN switch to read the same Radar in Czech or English. English content is published after its translation has been checked, so new and older items may appear later.
Everything you need to navigate the AI world
Today’s briefingThe “What is worth attention” selection sits beside Live · AI Flash, followed by research and links to other Radar sections. On mobile, these blocks appear one below another.
AI FlashAn ongoing feed of brief updates with an evidence status. Links lead to a Radar detail page when one is ready, otherwise to the original source. You can also find reset and outage histories here.
Practical applicationsWhat new tools and features can do, what you can try and what their actual impact could be.
Model selectionModel comparison by type of work, capabilities, price and speed.
Research and archiveA separate research overview, topic search and older events by date.
One event, everything that matters
Each row represents one event — not one article. At a glance, you can see its significance, credibility and main point.
Illustrative example, not a current news item.
Importance: ▮▮▮ majorOpenAIModels✓ 6
Agent mode is available to all paying users
Until now, the mode was available only on the highest plan; it is now available on all paid tiers without a waitlist.
▮▮▮ major · ▮▮ important · ▮ we're tracking = how significant the change is✓ 6 = six independent publishers, not the number of articles or feeds✓ official = a clear release, law or incident is substantiated by the relevant authority1 source = no independent confirmation yetbold = who is behind the changegray text = a brief summary of what happened
The detail page contains a fuller summary, its significance and original sources. Practical impact appears in the detail and the For individuals and For businesses views. An AI Flash item reaches the main selection only after it has been expanded and meets the publication rules.
The same news, two practical uses
We first summarize each event in the same way for everyone. Based on those same facts, we then explain what the change means for your own use and what it could mean for how a company operates.
For individualsWhat you can use or try, how the change can help you at work and what to watch out for.
For businessesWhat impact the change could have on processes, costs, risks and other business decisions.
Today’s briefing is the same for everyone. Pages
For individuals and For businesses
can be found in the main navigation — they select only events relevant to the given use case.