Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

846 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

A monotonic framework for evaluating operational risk in wildfire firefighting operations

The study proposes a new framework for evaluating wildfire risk prediction systems that, instead of traditional metrics, measures whether increasing risk scores consistently correspond to greater operational demands. A comparison of the expert-based DFE index, GRU models and FARS (a hybrid system with an LLM) in French…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

What red-team evaluations of AI models can and cannot prove

Researchers define a mathematical "evidential ceiling" framework for measuring what red-team benchmarks can establish about AI safety under a limited budget. They found that current tests are adequate for common harms, but millions of times insufficient for rare catastrophic incidents.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

A hybrid system for predicting roadblocks in Bolivia

A research paper combines time series modeling (Prophet) with NLP on a 6-year corpus of Bolivian newspapers. The system achieves an AUC-ROC of 0.677 for one-day forecasting and reduces Brier Score by 10.9% compared with statistical models, capturing signals of social tension in media discourse to predict road…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

MissionBench: A benchmark for evaluating multimodal models in autonomous flight missions

MissionBench is introduced with 120 missions in 5 simulated 3D environments for evaluating 22 multimodal models. The best model achieved a success rate below 35 percent (humans: 84.4 percent), illustrating the complexity of multilevel embodied tasks. Results show that larger models have better capabilities…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Physiological signals for detecting talking-face deepfakes

A research team proposed a detection system for talking-face deepfakes that synthesize video from a photograph and audio. The method extracts physiological signals (rPPG) through RhythmFormer and trains lightweight classifiers. On Celeb-DF++, a 1D ResNet achieved an AUC of 0.806; performance varies considerably…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

A discrete action space as a prerequisite for GRPO convergence in continuous control

The research tests Group Relative Policy Optimization (GRPO) on quadcopter control using the Qwen-0.5B model. Without modifications, the model collapses to a zero action (0% success), but replacing the continuous action space with a 5-way choice of PID presets achieves 98.6–100% success. The classical PID method achieves…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Context anxiety in large language models: Why they fail despite sufficient capabilities

The scientific study describes the phenomenon of ‘context anxiety’: reasoning models have the capabilities to solve problems but fail because of premature self-doubt and poor estimates of the tokens needed. The authors show that models can learn better strategies without this problem, without having to grow in size.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Developing and validating a Spanish version of the psychological dependence on LLMs scale

A study validates the Spanish version of a psychological dependence on large language models scale (LLM-D12-SP). A sample of 386 participants confirmed a two-factor structure: instrumental dependence (using LLMs for tasks and decisions) and relational dependence (psychological attachment). Cronbach's…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

A factorial study of synthetic data generation for machine translation of low-resource languages using grammar reference books

The research team used LLMs to extract grammatical rules and examples from grammar reference books and generated synthetic parallel corpora for fine-tuning translation models. Testing on three endangered languages (Kalamang, Tuatschin, Mandan) showed an improvement of +8.8 to +3.3 ChrF++ points. The study maps…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

J-CoT: Chain-of-Thought in J-Space

A research approach combining chain-of-thought with latent representation. J-CoT proposes an intermediate representation based on vocabulary-indexed coefficients instead of fully verbalizing each step. In tests, it outperformed existing latent-reasoning methods on mathematical, scientific, and coding tasks.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Evaluation design affects conclusions about MeSH terms in medical classification

A study examines how evaluation design affects comparisons of expert-assigned and automatic MeSH terms for classifying medical abstracts. On the Statins topic, the difference between approaches (WSS@95%) changes from +0.096 in 5-fold evaluation to +0.021 in 10-fold evaluation. In the 10-fold design, BiomedBERT and bag-of-words produce similar…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Benchmarking fine-tuning and retrieval strategies on the NRC Reactor Operator examination

A study evaluates Gemma 4 31B-IT on 14 NRC Reactor Operator examinations from 2015–2021. Comparing 8 fine-tuning and retrieval-augmented generation configurations showed that SFT with fixed-size chunking RAG passed 8 of 14 examinations (79.7 % accuracy). RAFT lagged behind SFT; the choice of chunking strategy…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

A semiotic logical hexagon for logical reasoning in LLMs

A research paper proposes HexLogicAgent, a framework that improves logical reasoning in LLMs by organizing semantics. The key finding: failures arise more from weak semantic representations than from deductive logic. Logical hexagon theory models the complete structure of semantic oppositions, which slows…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

SceneActBench: Evaluating agents' ability to act in 3D environments

Researchers introduced the SceneActBench benchmark for testing vision-language model agents capable of acting in 3D scenes. The benchmark contains five tasks with 520 cases. Testing eleven proprietary VLM configurations produced scores of 38.6–50.2 without consistent performance.

arXiv cs.AI (Artificial Intelligence) Original source ↗