Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

825 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Predicting Himalayan glacial floods: testing satellite data and ML models using freely available data

The research study tests predictions of glacial lake events and landslides in the Himalayas using freely available satellite data (radar interferometry, weather) and machine learning models. Using data on 589 lake outbursts and thousands of landslides, the authors compare deep learning with gradient boosting. Antecedent weather…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

More principled knowledge editing methods for LLM reasoning

A perspective in Nature Machine Intelligence points out that knowledge in LLMs forms an interconnected system. Updating knowledge without considering its interdependencies disrupts a model's logical reasoning ability. Three directions are proposed: editing deductive closure, integrating the model's beliefs and…

Nature Machine Intelligence Original source ↗
Research only one source so far

SAG: structured retrieval-augmented generation with dynamic hyperedges

The research paper proposes SAG, a new approach to retrieval-augmented generation (RAG). SAG indexes documents as event-entity pairs rather than traditional knowledge graphs. It achieves the best benchmark results, with 80.36% Recall@5 on MuSiQue (11.52 points better than competing approaches).

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

BEST-KAG: knowledge graphs and LLMs for questions about building standards

A research team developed BEST-KAG, a system for answering questions about building standards. It combines multimodal knowledge graphs and LLMs. The system processed 251 building standards with 171,652 nodes and 310,914 edges. In tests, it achieves improvements of up to 74% over baseline LLMs.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

FarSky: a generative model for intra-hour solar irradiance forecasting

The scientific preprint introduces FarSky, a method for forecasting solar irradiance using a latent diffusion model over a task-aware representation. It improves accuracy by up to 11 percentage points and achieves an F1-score above 60% in ramp event detection. Tested on data from Almería, Spain.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

LoongReflect: improving reflection in long-horizon agents through global distillation

A research team introduced LoongReflect, a method for training LLM agents with improved reflection in long-term decision-making. It combines knowledge distillation from a privileged teacher with GRPO trajectory optimization. Tests on retrieval with RAG and mathematical reasoning show consistent improvements over…

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

From prompting to behavioral alignment: personalized LLM judges for evaluating recommender systems

The research paper addresses an LLM failure (bidirectional rationalization) in recommender system evaluation. Behavioral alignment with fine-tuning achieves a 32.19% improvement in Macro-F1 over zero-shot. The approach matches the production baseline without manual overhead and offers interpretable reasoning.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

Group Relative Policy Optimization for financial advice: improvements over commercial models in a causal audit

A research team fine-tuned an open language model using Group Relative Policy Optimization (GRPO) with an LLM-as-a-judge reward function to generate financial recommendations. In a causal audit with a CATE estimator, the model achieves approximately 2× higher gross-profit lift ($0.0228 versus $0.0104). It has the lowest…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Multilingual API calling: supervised training more effective than RL

The research addresses Argument Language Mismatch – a phenomenon in which models call the correct tool but generate arguments in the wrong language. Supervised training (SFT) serves as a strong baseline with performance comparable to RL; methods such as GRPO bring only marginal improvements. The authors verified that careful…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Comparing the trustworthiness of small language models: pretrained versus compressed models

The scientific study evaluates the trustworthiness of small language models (SLMs) across fairness, robustness, privacy and ethics. Quantization preserves trustworthiness better than pruning. Compressing trustworthy models through quantization produces more reliable SLMs than training from scratch.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

EnterpriseRAG: a benchmark for testing LLM instruction following and robustness under non-ideal enterprise retrieval conditions

The research revealed a critical gap in enterprise RAG: LLMs satisfy 80% of individual constraints, but only 26.8% of responses satisfy all of them simultaneously. The new EnterpriseRAG benchmark, with 983 expert-validated samples across 6 domains, tests 13 SOTA models and simulates three critical failure modes: retrieval noise…

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

More efficient unlearning: identifying data with low influence reduces computational costs

Research from Apple and Harvard proposes an unlearning method that identifies training data with negligible influence on the model and excludes it from the removal process. The approach theoretically saves up to 50 % of computational costs. It uses influence functions for analysis on both language and image tasks.

Apple Machine Learning Research Original source ↗
Research only one source so far

Survey: companies lack data for AI agents

A survey of 300 leaders showed that AI has access to only 45% of corporate data on average (just 30% at lagging companies). Companies with 70%+ access trust agents 100%, while others struggle with scaling and speed. Gartner predicted that by 2027, agents will augment half of business…

MIT Technology Review — AI section Original source ↗
Research only one source so far

MindTopo benchmark reveals VLM limitations in topological reasoning

Microsoft Research introduced MindTopo, a benchmark for evaluating VLMs' ability to understand topological properties (connectivity, enclosure, knotting). Models perform better at static recognition but fail at planning tasks where they must maintain their understanding throughout actions. All models remain below…

Microsoft Research Blog Original source ↗