barga and Julia-1, compared on identical held-out questions. The earlier checkpoint audit follows below.
Stored internal evaluations, not independently rerun. Both checkpoints are shown as equal means of training seeds 17 and 18. These are question accuracies, not game win rates.
The five game evaluations use paired English (EN) and Nepali (NE) views of each underlying case. Language pairs and repeated training seeds are identified below so the question counts are not mistaken for independent game matches.
barga and Julia-1 on identical questions
Stored report measurements
Benchmark
Questions
barga
Julia-1
Calls · English
2,589
1,993 · 77.0%
580 · 22.4%
Calls · Nepali
2,589
1,997 · 77.1%
306 · 11.8%
Typed decisions
2,000
1,527 · 76.35%
1,451 · 72.55%
MNLI-Nepali
300
193 · 64.3%
93 · 31.0%
Kura-style calls
2,589
1,867 · 72.1%
—
Dhumbal
1,440
893 · 62.0%
Not trained
Ludo
1,592
1,001 · 62.9%
Not trained
Marriage
1,494
713 · 47.7%
Not trained
Call Break
1,828
873 · 47.8%
Not trained
Open-Jev · Julia-1 leads
9,172
57.6%
87.8%
Hard cases · Julia-1 leads
1,500
395 · 26.3%
505 · 33.7%
Emotion · Julia-1 leads
100
81 · 81.0%
86 · 86.0%
Full counts and caveats on later pages.
Devices and open weights
Stored report measurements
Device
Time per turn
Source
Apple M3 Pro CPU
0.87 s
CPU simulator report
T4 GPU
41 ms
Live service report
140.6M parameters · Apache-2.0 · Hugging Face
CPU: end-to-end time. GPU: model time; bgc1 measured separately.
Stored held-out evaluations at the source revision below, not independently rerun. barga and the earlier model are equal means of training seeds 17 and 18. Julia-1 answered the identical questions through its own runtime; questions it could not accept count as wrong.
Typed decisions: barga training seeds 17: 1527 / 2000 (76.350000%); 18: 1515 / 2000 (75.750000%). The paper labels seed 17; both seeds are disclosed here.
Marriage: barga 47.556894%. 1,494 held-out questions · EN + NE. barga correct: 713 / 708 of 1494. Training seeds 17 / 18. 747 paired test cases. Discard-choice accuracy regressed and remained below the option-position baseline. Overall question accuracy is not playing strength. Julia-1 has no adapter for this game.
Dhumbal: barga 62.083333%. 1,440 held-out questions · EN + NE. barga correct: 893 / 895 of 1440. Training seeds 17 / 18. 720 paired test cases. Synthetic, rule-checked positions with heuristic and rollout teachers; the Nepali templates have not had native review. Julia-1 has no adapter for this game.
Ludo: barga 62.688442%. 1,592 held-out questions · EN + NE. barga correct: 1001 / 995 of 1592. Training seeds 17 / 18. 796 paired test cases. Rule and strategy questions from synthetic positions. These scores do not measure full-game win rates. Julia-1 has no adapter for this game.
Chess: barga 41.278375%. 1,674 held-out questions · EN + NE. barga correct: 716 / 666 of 1674. Training seeds 17 / 18. 837 paired test cases. Binary move-legality accuracy is near chance. Legal moves are always constrained by a separate rules engine. Julia-1 has no adapter for this game.
Call Break: barga 46.663020%. 1,828 held-out questions · EN + NE. barga correct: 873 / 833 of 1828. Training seeds 17 / 18. 914 paired test cases. Card-play question accuracy regressed and remained below the option-position baseline. The implemented variant is must-trump v2. Julia-1 has no adapter for this game.
Training and size
Trained on how Nepal actually talks: a small open model that reads a situation in Nepali or English and picks one of the options you give it.
Trained on real Nepali and English customer-service call decisions under an agreement with a partner.
Without real-call training, real-call accuracy drops from 77% to 45%.
Questions in Nepali or English. Robust to Nepali speech-recognition errors. Nepali games.
140M parameters. Under a second per turn on a laptop CPU. Open under Apache-2.0 for use in-house.
barga is a beginner at chess.
stronger as the tiger than as the goats.
Probabilities are barga’s preference among the options shown, not win chances.
The original checkpoint names below identify the historical reports.
खेलको मापन
Marriage
1,494 held-out questions · EN + NE
1,494 held-out questions per seed: 747 English + 747 Nepali, representing 747 paired test cases.
Marriage: held-out question accuracy, two training seeds
Checkpoint
Two-seed mean
Seed 17: correct
Seed 18: correct
barga
47.556894%
713 / 1,494
708 / 1,494
earlier model
31.793842%
483 / 1,494
467 / 1,494
Stored reports · two-seed means. Paired language views. Not a game win rate.
barga correct: 713 / 708 of 1494; earlier model correct: 483 / 467 of 1494. Training seeds 17 / 18. 747 paired test cases.
What to keep in mind. Discard-choice accuracy regressed and remained below the option-position baseline. Overall question accuracy is not playing strength.
खेलको मापन
Dhumbal
1,440 held-out questions · EN + NE
1,440 held-out questions per seed: 720 English + 720 Nepali, representing 720 paired test cases.
Dhumbal: held-out question accuracy, two training seeds
Checkpoint
Two-seed mean
Seed 17: correct
Seed 18: correct
barga
62.083333%
893 / 1,440
895 / 1,440
earlier model
38.680556%
563 / 1,440
551 / 1,440
Stored reports · two-seed means. Paired language views. Not a game win rate.
barga correct: 893 / 895 of 1440; earlier model correct: 563 / 551 of 1440. Training seeds 17 / 18. 720 paired test cases.
What to keep in mind. Synthetic, rule-checked positions with heuristic and rollout teachers; the Nepali templates have not had native review.
खेलको मापन
Ludo
1,592 held-out questions · EN + NE
1,592 held-out questions per seed: 796 English + 796 Nepali, representing 796 paired test cases.
Ludo: held-out question accuracy, two training seeds
Checkpoint
Two-seed mean
Seed 17: correct
Seed 18: correct
barga
62.688442%
1,001 / 1,592
995 / 1,592
earlier model
39.949749%
637 / 1,592
635 / 1,592
Stored reports · two-seed means. Paired language views. Not a game win rate.
barga correct: 1001 / 995 of 1592; earlier model correct: 637 / 635 of 1592. Training seeds 17 / 18. 796 paired test cases.
What to keep in mind. Rule and strategy questions from synthetic positions. These scores do not measure full-game win rates.
खेलको मापन
Chess
1,674 held-out questions · EN + NE
1,674 held-out questions per seed: 837 English + 837 Nepali, representing 837 paired test cases.
Chess: held-out question accuracy, two training seeds
Checkpoint
Two-seed mean
Seed 17: correct
Seed 18: correct
barga
41.278375%
716 / 1,674
666 / 1,674
earlier model
29.121864%
506 / 1,674
469 / 1,674
Stored reports · two-seed means. Paired language views. Not a game win rate.
barga correct: 716 / 666 of 1674; earlier model correct: 506 / 469 of 1674. Training seeds 17 / 18. 837 paired test cases.
What to keep in mind. Binary move-legality accuracy is near chance. Legal moves are always constrained by a separate rules engine.
खेलको मापन
Call Break
1,828 held-out questions · EN + NE
1,828 held-out questions per seed: 914 English + 914 Nepali, representing 914 paired test cases.
Call Break: held-out question accuracy, two training seeds
Checkpoint
Two-seed mean
Seed 17: correct
Seed 18: correct
barga
46.663020%
873 / 1,828
833 / 1,828
earlier model
28.528446%
530 / 1,828
513 / 1,828
Stored reports · two-seed means. Paired language views. Not a game win rate.
barga correct: 873 / 833 of 1828; earlier model correct: 530 / 513 of 1828. Training seeds 17 / 18. 914 paired test cases.
What to keep in mind. Card-play question accuracy regressed and remained below the option-position baseline. The implemented variant is must-trump v2.
निर्णयको मापन
Typed decisions
2,000 held-out questions
2,000 held-out English questions per seed.
Typed decisions: held-out question accuracy, two training seeds
Checkpoint
Two-seed mean
Seed 17: correct
Seed 18: correct
barga
76.050000%
1,527 / 2,000
1,515 / 2,000
earlier model
76.025000%
1,518 / 2,000
1,523 / 2,000
Stored reports · two-seed means. Training split used; test excluded.
barga seed 17: 76.35%; seed 18: 75.75%. earlier model seed 17: 75.90%; seed 18: 76.15%. The mean difference is 0.025 percentage points.
What to keep in mind. This is task-specific test accuracy, not zero-shot or broad reasoning performance.
नेपालीमा संवाद
Nepali call decisions
Nepali transcripts · English questions
2,589 held-out questions per seed across 94 held-out calls. Transcript states include Nepali and Nepali–English mixtures; the questions are in English.
Nepali call decisions: held-out question accuracy, two training seeds
Checkpoint
Two-seed mean
Seed 17: correct
Seed 18: correct
barga
76.554654%
1,993 / 2,589
1,971 / 2,589
earlier model
76.979529%
2,006 / 2,589
1,980 / 2,589
Stored reports · two-seed means. Teacher labels, not human gold.
barga mean: 76.554654%; earlier model mean: 76.979529%. The newer checkpoint is lower on this dataset.
What to keep in mind. Measures fit to this call dataset. No real-call text is included. A native-reviewed Nepali benchmark remains missing.
मापनको सन्दर्भ
Read the fine print
Every result has a context.
These are stored question-level evaluations, not independently rerun matches.
A stronger overall score can hide weaker decision families. Legal actions come from separate rule engines.
These are internal held-out question-level accuracies. They are not game win rates or evidence of full-game playing strength.
Both checkpoints use an equal mean of training seeds 17 and 18. Each seed answers the same test questions; two seeds do not double the number of independent test cases.
For each game, the English and Nepali questions are paired views of the same underlying cases. They are not independent samples.
The mean is 100 × (correct in seed 17 + correct in seed 18) ÷ (2 × questions per seed). Percentages are displayed to six decimal places; exact counts are shown above.
Typed-decisions training data was used. The result is task-specific, not zero-shot or a broad reasoning benchmark.
Call labels are teacher proposals, not human gold. No real support-call text is included here. A native-reviewed Nepali benchmark remains missing.
No general superiority, calibrated confidence or browser latency is established by these reports.
barga / earlier model · two training seeds · snapshot 64133e0
The room’s saved synthetic call example is barga, training seed 17, recorded on CPU with four threads on 2026-10-05. It is saved output, not live inference.