Skip to content
worth noting New models

The Mogan AI team released the Turkish encoder model MoganBert-TR with the CLM-to-MLM training procedure

only one source so far

The Mogan AI team released the 149M-parameter Turkish encoder model MoganBert-TR, trained on 237.3 billion tokens using the CLM-to-MLM procedure; it achieves 78.41 points on TrGLUE, the best result among Turkish ModernBERT models. The accompanying smaller embedding model achieves 99.5 % of the performance of the teacher.

The Mogan AI team published MoganBert-TR, a 149M-parameter Turkish encoder model trained from scratch on a language-filtered corpus. The model was trained on 237.3 billion tokens using a two-stage CLM-to-MLM procedure: first causal language modeling, then masked language modeling for the remainder of training, with the transition between stages occurring within the stable phase of the WSD schedule.

According to the company, in a controlled ablation test with the same step budget, this procedure outperforms pure MLM by 2.7 to 3.7 times on the Turkish retrieval benchmark MS MARCO. The identified mechanism is embedding geometry — with pure MLM, a single direction accounts for 28.1 % of the variance, compared with only 11.9 % with the proposed procedure. Long-context extension and learning rate decay were subsequently split into two branches: completing the final part of the decay at a context length of 1024 improved the TrGLUE average by 0.49 ± 0.26 points across five seed pairs (p = 0.013) and outperformed a model-soup alternative by 0.75 points at the cost of approximately 4.3 % in additional costs.

MoganBert-TR achieves 78.41 points on TrGLUE, which the company says is the best result among the Turkish ModernBERT models compared, and 77.73 points on TabiBench, where it leads in two of eight categories, with the largest margin in code search (+3.62 points over TabiBERT). The derived embedding model MoganBert-Embed, created through distillation from a teacher and contrastive fine-tuning with multiple signals, ranked first among student models on MTEB(Turkish) with a score of 68.30, corresponding to 99.5 % of the performance of its 7.57B-parameter teacher with a backbone 51 times smaller. According to the company, the accompanying tokenizer with 50 048 tokens outperforms all the Turkish tokenizers compared in compression and fertility on two independent test sets.

The model weights, tokenizer, embedding model and evaluation code are publicly available on HuggingFace under the moganai account.

What changed

Why it matters

This is a freely available foundation model trained specifically for Turkish, with public weights, a tokenizer and evaluation code, so developers do not need to train their own encoder from scratch. According to the company, the accompanying embedding model MoganBert-Embed also achieves almost the same quality as a significantly larger teacher at a size 51 times smaller, which is relevant to deploying search and embedding systems for Turkish with lower computational requirements.

Release card

MoganBert-TR

Mogan AI

open weights
Specifications
149M
Availability
Weights, tokenizer, embedding model and evaluation code publicly available on HuggingFace under the moganai account.
Documented measurements
  • TrGLUE 78,41 bodu According to the source, the best result among the Turkish ModernBERT models compared.
  • Turecké MS MARCO (retrieval, kontrolovaná ablace) 2,7-3,7krát vyšší skóre oproti čistému MLM tréninku při stejném rozpočtu kroků The CLM-to-MLM training procedure improves retrieval task results compared with MLM-only training under comparable conditions.
According to the sources, it is suitable for
  • Turkish text search (retrieval)
  • code search in Turkish
  • generating embeddings using the derived MoganBert-Embed model

The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →

Two audiences, two different impacts

What this means

01

For individuals

Developers and researchers working with Turkish text gain a freely available encoder model and embedding model that they can directly integrate into their own NLP tasks (search, classification, embeddings) without having to train their own model from scratch.

What to do Try MoganBert-TR and MoganBert-Embed from HuggingFace for your own projects that process Turkish text.
More practical updates →
02

For a business

According to the Mogan AI team, companies running search or embeddings for Turkish text can deploy a significantly smaller model (a backbone 51x smaller than the teacher) while retaining 99.5 % of its performance on MTEB(Turkish), reducing inference and infrastructure costs.

Development
What to decide Consider replacing the existing embedding model for Turkish with MoganBert-Embed and verify the benefits on your own data, because according to the company, it achieves 99.5 % of the performance of a significantly larger teacher.
More business impacts →
Curriculum learning embedding model encoder MoganBert-TR TrGLUE Turkish

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
arXiv cs.CL (Computation and Language / NLP) research source · first detected MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM-to-MLM Curriculum