Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

230 published research events 4 new research papers today

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Lingtai study tracks uncertainty during LLM inference, but does not reliably determine correctness

The study presents Lingtai, a layer for tracking concept signals during LLM inference without training probes. The signals relate to predictive uncertainty, but in tests they do not provide a stable indicator of correctness. According to the authors, the overhead during code generation is 0.7–1.6% per token.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Study compares evaluation methods for Romanian RAG systems using LLMs

The study adapted the Ragas framework for Romanian and created the AdminRo-Eval dataset. The authors report 96% agreement with human evaluation for Faithfulness when decomposing the evaluation with the Gemini 2.5 Pro model, and 90% for Answer Relevance when comparing answers.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Study uses Fuzzy Cognitive Maps to generate synthetic health data

The study presents the generation of synthetic health data using Fuzzy Cognitive Maps running exclusively on CPU. On the Heart Disease dataset, the authors report an accuracy of up to 0.81 and AUROC of up to 0.90; the results are compared with the TVAE and Gaussian Copula methods.

arXiv cs.LG (Machine Learning) Original source ↗
Research only one source so far

Research survey classifies RAG methods along four axes

The study organizes RAG research according to retrieval efficiency, robustness and safety, interactive procedures, and multi-step reasoning. It compares methods, architectures, and evaluation approaches and describes persistent reliability and scaling challenges.

arXiv cs.AI (Artificial Intelligence) Original source ↗
Research only one source so far

CypherTurn study reveals low reliability of models in multi-step graph database querying

The study introduces CypherTurn with 721 conversations and 5 927 steps over 7 knowledge graphs. In evaluating 15 models, the best model achieved 64.7% query execution correctness; full-conversation correctness remained below 5%. A higher action budget did not resolve the gap in autonomous operation.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Study: probe-guided fine-tuning improves model safety and preserves their controllability

According to the authors, continuously updated probes during training reduce harmfulness and improve the truthfulness of models while preserving usefulness. The method shows a better safety-to-usefulness ratio than DPO and inference-time steering, and it preserves the ability to inspect internal representations.

arXiv cs.LG (Machine Learning) Original source ↗