Nirnoy নির্ণয়

Benchmarks · updated 9 October 2026

How Nirnoy compares on Bangla and Banglish

Nirnoy comes in two sizes: Nirnoy-flash, small and fast, and Nirnoy 12B, the most accurate on harder tasks. We asked both and five other models the same questions, with the same answer options, mostly on human-labelled Bangla and Banglish data. Scores are macro-F1 (%), higher is better; ★ marks the best model in each category, and any within 0.1 points of it, which count as tied for first: the same hosted model moved that much between two runs (Jev scored 37.4 on emotion on 6 October 2026 and 37.5 on emotion on 9 October 2026).

9 of 9categories where a Nirnoy model is firstNirnoy 12B 6, Nirnoy-flash 3; Jev 1
90.1Nirnoy 12B on SMS scam detectionNirnoy-flash 82.8, Jev 88.4, Clef 85.3
77.0Nirnoy-flash on everyday tasksNirnoy 12B 76.5, Jev 70.1, Clef 66.9
76.4Nirnoy-flash on Banglish everyday tasksNirnoy 12B 76.3, Jev 61.0, Clef 54.4

At a glance

CategoryNirnoy-flashNirnoy 12BJevClefClef-flashGemma 12BGemma E4B
Everyday tasks
Assistant commands: intent85.6 ★82.876.679.879.5––
Assistant commands: scenario92.3 ★90.884.386.988.5––
Hate speech in Bangla and Banglish84.1 ★82.570.261.658.1––
Code-mixed sentiment76.277.1 ★63.854.751.5––
News topic88.188.3 ★88.184.379.7––
Emotion36.137.4 Tied for first: 37.4 is within 0.1 points of the best score, 37.5. The same hosted model, run again on another day, moves this much: Jev scored 37.4 on emotion on 6 October 2026 and 37.5 on emotion on 9 October 2026. So gaps this small count as a tie.37.5 ★34.131.6––
Scam detection
SMS scam detection82.890.1 ★88.485.383.888.683.1
Scam messages, private set86.887.3 ★–67.359.373.082.6
Decision scenarios, synthetic (accuracy)
Decision scenarios77.280.1 ★75.774.667.770.867.8

Everyday tasks

A human-labelled test of 11,063 questions, every option shown. Nirnoy trained on other examples from the first four datasets; topic and emotion are new to it. The base Gemma models were not run on this test with every option shown.

Assistant commands: intent

Pick the right one of 60 intents · 1,500 commands

Data: MASSIVE (Bengali) (CC BY 4.0)

Example

কাল সকাল সাতটায় একটা অ্যালার্ম দিয়ে দাও

Which of 60 intents? → set an alarm

Nirnoy-flash
85.6 ★
Nirnoy 12B
82.8
Jev
76.6
Clef
79.8
Clef-flash
79.5
Gemma 12B
not run
Gemma E4B
not run

Assistant commands: scenario

Pick the right one of 18 areas · 1,500 commands

Data: MASSIVE (Bengali) (CC BY 4.0)

Example

আজ ঢাকায় বৃষ্টি হবে কি?

Which of 18 areas? → weather

Nirnoy-flash
92.3 ★
Nirnoy 12B
90.8
Jev
84.3
Clef
86.9
Clef-flash
88.5
Gemma 12B
not run
Gemma E4B
not run

Hate speech in Bangla and Banglish

Is a comment hateful? · 3,966 comments

Data: BanHate (MIT) and BanTH (MIT)

Example

tor moto faltu lok ar dekhini, chup thak

Is this hateful? → yes

Nirnoy-flash
84.1 ★
Nirnoy 12B
82.5
Jev
70.2
Clef
61.6
Clef-flash
58.1
Gemma 12B
not run
Gemma E4B
not run

Code-mixed sentiment

Banglish comments, including mixed feelings · 2,024 comments

Data: BnSentMix (MIT)

Example

Movie ta overall bhalo, but ending ta ektu boring laglo.

Positive, negative, neutral or mixed? → mixed

Nirnoy-flash
76.2
Nirnoy 12B
77.1 ★
Jev
63.8
Clef
54.7
Clef-flash
51.5
Gemma 12B
not run
Gemma E4B
not run

News topic

Pick one of 7 topics (not in Nirnoy's training) · 204 headlines

Data: SIB-200 (Bengali) (CC BY-SA 4.0)

Example

বিশ্বকাপ বাছাইপর্বে শেষ মিনিটের গোলে জিতল বাংলাদেশ

Which topic? → sports

Nirnoy-flash
88.1
Nirnoy 12B
88.3 ★
Jev
88.1
Clef
84.3
Clef-flash
79.7
Gemma 12B
not run
Gemma E4B
not run

Emotion

Pick one of 6 emotions (not in Nirnoy's training) · 1,869 texts

Data: EmoNoBa (CC BY 4.0)

Example

রেজাল্ট দেখে চোখে পানি চলে এলো, এত খুশি আগে কখনো লাগেনি!

Which emotion? → joy

Nirnoy-flash
36.1
Nirnoy 12B
37.4 Tied for first: 37.4 is within 0.1 points of the best score, 37.5. The same hosted model, run again on another day, moves this much: Jev scored 37.4 on emotion on 6 October 2026 and 37.5 on emotion on 9 October 2026. So gaps this small count as a tie.
Jev
37.5 ★
Clef
34.1
Clef-flash
31.6
Gemma 12B
not run
Gemma E4B
not run

Scam detection

Real messages Nirnoy never trained on: a public set of Bengali SMS (scored on a held-back half, once) and a private set of messages. Jev was not run on the private set, because it is never sent to outside services; with 110 messages, gaps under about 8 points are within noise.

SMS scam detection

Public set of real Bengali SMS (MIT licence): scam, spam or legitimate · 701 messages

Data: Bengali SMS smishing dataset (MIT)

Example

অভিনন্দন! আপনার নম্বর ৫০,০০০ টাকা জিতেছে। টাকা পেতে এখনই এই লিংকে ঢুকে আপনার মোবাইল ব্যাংকিং পিন দিন।

Scam, spam or legitimate? → scam

Nirnoy-flash
82.8
Nirnoy 12B
90.1 ★
Jev
88.4
Clef
85.3
Clef-flash
83.8
Gemma 12B
88.6
Gemma E4B
83.1

Scam messages, private set

Real messages, with the question in Bangla and in English · 110 messages

Data: A private set of real messages; not published.

Example

আপনার মোবাইল ব্যাংকিং একাউন্ট আজ বন্ধ হয়ে যাবে। চালু রাখতে OTP কোডটি এই নম্বরে পাঠান।

এই মেসেজটি কি প্রতারণা, স্প্যাম নাকি বৈধ? / Scam, spam or legitimate? → scam

Nirnoy-flash
86.8
Nirnoy 12B
87.3 ★
Jev
not run
Clef
67.3
Clef-flash
59.3
Gemma 12B
73.0
Gemma E4B
82.6

Synthetic decision scenarios

10,000 everyday decisions of ten kinds, written in Bangla by a language model, which also set the answers. Useful for breadth, but the answers are not human-checked.

Decision scenarios

Routing, safety checks, answer checks, record matching and six more kinds of decision; accuracy (%) · 10,000 rows

Data: A private synthetic set written with a language model; not published.

Example

আমার অর্ডার এখনো আসেনি, টাকা ফেরত চাই।

Route to: sales, refunds, technical support or delivery? → refunds

Nirnoy-flash
77.2
Nirnoy 12B
80.1 ★
Jev
75.7
Clef
74.6
Clef-flash
67.7
Gemma 12B
70.8
Gemma E4B
67.8

How we measured

Same questions for every model. Each model saw the same text, question and full list of answer options. Jev, Clef and Clef-flash answered through their public APIs on OpenRouter (October 2026). The private scam messages never leave our machines, so there Clef and Clef-flash ran from their released weights on our GPUs and Jev was not run. The Gemma 4 base models are untrained; Nirnoy-flash is built on Gemma 4 E4B and Nirnoy 12B on Gemma 4 12B.

Macro-F1 averages the F1 score over the answer classes, so a model cannot score well by always giving the most common answer.

Held-back half. The public SMS set is split in two by record. We tuned only on one half; the score here is from the other half, scored once.

What Nirnoy saw in training. Neither model was trained on the scam sets or the decision scenarios. For the everyday tasks they were trained on other examples from the intent, scenario, hate and sentiment datasets; topic and emotion are new to them.

Examples on this page are written for illustration and are not taken from the test sets.