Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

825 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

CrisisKD: five-stage knowledge distillation for sentiment and emotion analysis in crisis discourse

The CrisisKD paper describes a framework for sentiment and emotion analysis on social media using knowledge distillation. The author releases a dataset with 50 615 aspect labels and scripts as open-source. The Qwen2.5-7B model achieves an improvement of 7.9 F1 points in aspect extraction and 17 points in emotion accuracy.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

ModularPhaseNet: discrete phase geometry for transformers

The authors propose ModularPhaseNet, which discretizes complex phase geometry into cyclic groups modulo a prime number. The method extends standard transformers with an auxiliary phase channel without quantum hardware. The aim is to improve semantic hierarchy, contextual alignment, contradiction detection and prediction…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Framework La Agente 'Optima for self-driving laboratories with AI agents

The new framework La Agente 'Optima combines LLM agents with Bayesian optimization for automated scientific experiments. In trials, it increased yield from 30 % to 59 % (23 experiments) and adjusted parameters in optimization campaigns. The system requires less time and fewer materials than human control.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

PetQA: A Benchmark for Evaluating Veterinary Knowledge and Clinical Decision-Making in LLM Models

The new Korean benchmark PetQA for evaluating veterinary knowledge in LLM and LVLM models contains 10 076 text-based and 8 751 multimodal QA pairs provided by veterinarians. The study evaluates 18 models using ROUGE, BERTScore and LLM-as-judge metrics with three methods: zero-shot inference, retrieval-augmented generation and…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Artificial Analysis updated the Intelligence Index after the GPT-6 Astra score drew criticism

Artificial Analysis released version 4.2 of the Intelligence Index after its benchmarks failed to reflect the actual progress of GPT-6 Astra. The Astra model received a four-point increase, while the Claude Fable 5.1 model remains in first place. New benchmarks (AA-Briefcase and GDP.pdf) added, GPQA-Diamond removed. Private data…

The Decoder (daily AI news) Original source ↗
Research only one source so far

FLIWBO method for Bayesian optimization with adaptable input transformations

The research introduces FLIWBO, a Bayesian optimization method that selects input transformations from a finite library. It preserves convergence guarantees with √N complexity, improves sample efficiency under geometric mismatches (log-scaled parameters, localized peaks), and outperforms GP-UCB on hyperparameter…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

GPS-Bench: a benchmark for simulating governance policies with LLM agents

A research team introduced the GPS-Bench benchmark, which links policy proposals to relevant actors using public records (legislation, lobbying, corporate filings). It tests how multi-agent simulations with LLM models predict policy impacts. Fine-tuning on grounded data yielded the best…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

SVG-Score: a metric for evaluating SVG generation with a focus on alignment with human judgment

A research team introduced SVG-Score, an evaluation framework for text-to-SVG generation. It includes a dataset with human annotations and two evaluators focused on alignment with human judgment. It shows that the CLIP score responds poorly to errors made by SVG generators, such as incorrect colors and counts.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Localization of optimization value in the harness of LLM agents: an analysis of decomposed scaffolding

The HARNESSEVO method decomposes the textual scaffolding of frozen LLMs into four optimizable slots. Slot-level analysis on ALFWorld shows that almost all the optimization value (gain +0.119) lies in the reflection/control slot. The other slots are neutral. Uniform budget allocation across slots is…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

The CW-Net method explains decision-making in autonomous vehicles

MIT and the company Motional developed the CW-Net method, which translates decision-making by deep learning models in autonomous vehicles into understandable concepts. Testing showed that these explanations help drivers better predict vehicle behavior and improve safety. The research was published in Nature.

MIT News – Artificial intelligence Original source ↗
Research only one source so far

Relational Transformer embeddings in an LLM – study reveals the limits of the hybrid approach

A research paper tests inserting Relational Transformer embeddings into the Qwen 3.5-4B model. A study on 10 binary classification tasks across 6 relational databases shows that the hybrid model usually does not outperform a standalone Relational Transformer, is unstable and sensitive to data formats. The authors…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

DISTAL: Combining self-learning and distillation for structure-free prediction of material properties

DISTAL is a research approach for predicting material properties in low-data settings. It combines self-supervised pretraining on compositional data with knowledge distillation from a trained ALIGNN model. Across 39 benchmark tasks, it improved performance on 37 of them. The code and models will be released as open-source.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

The impact of retrieval, scoring and decoding on the performance and stability of LLM rerankers in recommender systems

A study on the ReDial dataset compared the performance of proprietary and open-weight LLM models as rerankers. Key finding: proprietary LLMs achieve NDCG@10 of 0.1497 compared with 0.0939 for collaborative filtering, but without strict candidate pool constraints, their advantage appears to be 0.2925. Open-weight models…

arXiv cs.CL (Computation and Language / NLP) Original source ↗