Skip to content

Research separate from news

What could become important next

A preprint is a signal, not a finished product or an independently confirmed result. We therefore track research papers separately and show the actual state of supporting evidence for each one.

846 published research events

Latest work

A significant claim from a single source is published only after further confirmation.

Research only one source so far

Study of diversity in LLM opinions: Persona conditioning more effective than simple methods

A scientific study tests interventions to increase diversity in LLM opinions. It found that simply increasing the temperature and providing instructions are ineffective; structured persona conditioning combined with interaction architectures is effective. The results are based on a factorial experiment with 100 questions and 7 models.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Measuring bias in LLM world models using prediction markets

Research shows that LLMs are biased by the text in the corpus they were trained on. The authors propose a methodology for measuring bias using prediction markets as an external reference. In an experiment on 111 Ukrainian markets (93 000 predictions, four models), they found that English-language news coverage systematically introduces bias…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

REGARD: Regional differences in the emotional framing of large language models

An arXiv study examines how 19 large language models frame political and cultural entities differently in the post-Soviet region. Using Valence-Arousal-Dominance profiling, it identified three model clusters with different affective behavior, with the rate of generic responses strongly correlated with lower…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

THOR: A hierarchical oscillatory framework for multi-hop question answering

The THOR research framework addresses two key problems in multi-hop question answering: attention decay during long reasoning chains and error accumulation. Inspired by theta-gamma oscillations in the brain, it separates global planning from local retrieval and introduces a verification mechanism to stop…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

From agent failures to textual policies: what works and what fails

Research results on TextGrad applied to agents. Handwritten policies improved the success of frozen 7B agents by 5.0 points, but policies generated from data performed no better than fixed prompting. The main challenge is reliably generating and selecting policies from experience, rather than executing them.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Tracing animacy circuits in large language models

Research maps how four open-weight LLMs process animacy (the distinction between living and non-living entities). Through circuit discovery and ablations, researchers identified causal mechanisms responsible for this behavior, but found that they are not strongly localized and generalize only partially across models.

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Skill-contracted agents for evidence-aware materials science literature analysis

A research paper introduces AlphaAgent, a framework for analyzing literature in materials science. It separates retrieval from report generation through explicit skill contracts. In a blind evaluation on 40 materials science questions, it substantially outperformed the baseline, especially in mechanistic…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Do diversity metrics really measure diversity? Auditing majority-voting gains in LLM ensembles

A study examines five diversity metrics in LLM ensembles across approximately 32 thousand subsets of 30 models. It finds that diversity measures model capability rather than diversity itself, and majority voting outperforms the best member in only ~10 % of cases. The results suggest limited usefulness of existing…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

How large language models assess writing quality — a research analysis

A study examined how DeepSeek and Qwen QwQ models evaluate literary quality. On a benchmark of 30 texts in 6 categories, they achieved 79.3 % accuracy. The models prefer intentionality, craftsmanship, depth and a distinctive voice. A second study showed that degrading structure and voice has a greater impact than simplifying lexical…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Natural language should not completely replace formal languages

An arXiv research paper argues that natural language should not completely replace formal languages for specification. The authors introduce task specificity and prove a specificity crossover theorem — there is a threshold beyond which expression in natural language is more demanding than direct formal…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Semantic field: a mathematical framework for modeling semantics and stable interpretation

A theoretical preprint paper formalizes the concept of Semantic Field Theory (SFT) for lexical semantics. It defines semantic fields as mathematical structures with deformations, interaction complexes, and Gaussian representations. It introduces energy minimization for stable interpretation. It is purely…

arXiv cs.CL (Computation and Language / NLP) Original source ↗
Research only one source so far

Detecting political evasiveness with the asymmetric verification method AsymVerify

An arXiv research paper presents AsymVerify, a methodology for classifying responses as clear, ambiguous, or refusals. The system achieved a Macro F1 score of 0.85 in SemEval-2026 Task 6 (second place among 41 teams), with asymmetric verification reducing errors at the boundary of ambiguous responses with…

arXiv cs.CL (Computation and Language / NLP) Original source ↗