Return to the room

barga · stored evaluation evidence

The field notes.

barga and Julia-1, compared on identical held-out questions. The earlier checkpoint audit follows below.

Stored internal evaluations, not independently rerun. Both checkpoints are shown as equal means of training seeds 17 and 18. These are question accuracies, not game win rates.

The five game evaluations use paired English (EN) and Nepali (NE) views of each underlying case. Language pairs and repeated training seeds are identified below so the question counts are not mistaken for independent game matches.

barga and Julia-1 on identical questions

Stored report measurements
BenchmarkQuestionsbargaJulia-1
Calls · English2,5891,993 · 77.0%580 · 22.4%
Calls · Nepali2,5891,997 · 77.1%306 · 11.8%
Typed decisions2,0001,527 · 76.35%1,451 · 72.55%
MNLI-Nepali300193 · 64.3%93 · 31.0%
Kura-style calls2,5891,867 · 72.1%—
Dhumbal1,440893 · 62.0%Not trained
Ludo1,5921,001 · 62.9%Not trained
Marriage1,494713 · 47.7%Not trained
Call Break1,828873 · 47.8%Not trained
Open-Jev · Julia-1 leads9,17257.6%87.8%
Hard cases · Julia-1 leads1,500395 · 26.3%505 · 33.7%
Emotion · Julia-1 leads10081 · 81.0%86 · 86.0%

Full counts and caveats on later pages.

Devices and open weights

Stored report measurements
DeviceTime per turnSource
Apple M3 Pro CPU0.87 sCPU simulator report
T4 GPU41 msLive service report

140.6M parameters · Apache-2.0 · Hugging Face

CPU: end-to-end time. GPU: model time; bgc1 measured separately.

CPU simulator report · Live service report

Stored held-out evaluations at the source revision below, not independently rerun. barga and the earlier model are equal means of training seeds 17 and 18. Julia-1 answered the identical questions through its own runtime; questions it could not accept count as wrong.

Typed decisions: barga training seeds 17: 1527 / 2000 (76.350000%); 18: 1515 / 2000 (75.750000%). The paper labels seed 17; both seeds are disclosed here.

Open-Jev macro average: barga 57.649055%; Julia-1 87.772647%. 12 families, 9,172 questions.

barga’s home tasks

Calls and reading comprehension: Julia-1 is a reference only. Task training and unsupported inputs affect this comparison.

Games: question accuracy, not playing strength

Training and size

Trained on how Nepal actually talks: a small open model that reads a situation in Nepali or English and picks one of the options you give it.

Trained on real Nepali and English customer-service call decisions under an agreement with a partner.

Without real-call training, real-call accuracy drops from 77% to 45%.

Questions in Nepali or English. Robust to Nepali speech-recognition errors. Nepali games.

140M parameters. Under a second per turn on a laptop CPU. Open under Apache-2.0 for use in-house.

barga is a beginner at chess.

stronger as the tiger than as the goats.

Probabilities are barga’s preference among the options shown, not win chances.

Stored reports at the source revision · Generated evidence data

Earlier checkpoint audit

The original checkpoint names below identify the historical reports.

खेलको मापन

Marriage

1,494 held-out questions · EN + NE

1,494 held-out questions per seed: 747 English + 747 Nepali, representing 747 paired test cases.

Marriage: held-out question accuracy, two training seeds
CheckpointTwo-seed meanSeed 17: correctSeed 18: correct
barga47.556894%713 / 1,494708 / 1,494
earlier model31.793842%483 / 1,494467 / 1,494

Stored reports · two-seed means. Paired language views. Not a game win rate.

barga correct: 713 / 708 of 1494; earlier model correct: 483 / 467 of 1494. Training seeds 17 / 18. 747 paired test cases.

What to keep in mind. Discard-choice accuracy regressed and remained below the option-position baseline. Overall question accuracy is not playing strength.

खेलको मापन

Dhumbal

1,440 held-out questions · EN + NE

1,440 held-out questions per seed: 720 English + 720 Nepali, representing 720 paired test cases.

Dhumbal: held-out question accuracy, two training seeds
CheckpointTwo-seed meanSeed 17: correctSeed 18: correct
barga62.083333%893 / 1,440895 / 1,440
earlier model38.680556%563 / 1,440551 / 1,440

Stored reports · two-seed means. Paired language views. Not a game win rate.

barga correct: 893 / 895 of 1440; earlier model correct: 563 / 551 of 1440. Training seeds 17 / 18. 720 paired test cases.

What to keep in mind. Synthetic, rule-checked positions with heuristic and rollout teachers; the Nepali templates have not had native review.

खेलको मापन

Ludo

1,592 held-out questions · EN + NE

1,592 held-out questions per seed: 796 English + 796 Nepali, representing 796 paired test cases.

Ludo: held-out question accuracy, two training seeds
CheckpointTwo-seed meanSeed 17: correctSeed 18: correct
barga62.688442%1,001 / 1,592995 / 1,592
earlier model39.949749%637 / 1,592635 / 1,592

Stored reports · two-seed means. Paired language views. Not a game win rate.

barga correct: 1001 / 995 of 1592; earlier model correct: 637 / 635 of 1592. Training seeds 17 / 18. 796 paired test cases.

What to keep in mind. Rule and strategy questions from synthetic positions. These scores do not measure full-game win rates.

खेलको मापन

Chess

1,674 held-out questions · EN + NE

1,674 held-out questions per seed: 837 English + 837 Nepali, representing 837 paired test cases.

Chess: held-out question accuracy, two training seeds
CheckpointTwo-seed meanSeed 17: correctSeed 18: correct
barga41.278375%716 / 1,674666 / 1,674
earlier model29.121864%506 / 1,674469 / 1,674

Stored reports · two-seed means. Paired language views. Not a game win rate.

barga correct: 716 / 666 of 1674; earlier model correct: 506 / 469 of 1674. Training seeds 17 / 18. 837 paired test cases.

What to keep in mind. Binary move-legality accuracy is near chance. Legal moves are always constrained by a separate rules engine.

खेलको मापन

Call Break

1,828 held-out questions · EN + NE

1,828 held-out questions per seed: 914 English + 914 Nepali, representing 914 paired test cases.

Call Break: held-out question accuracy, two training seeds
CheckpointTwo-seed meanSeed 17: correctSeed 18: correct
barga46.663020%873 / 1,828833 / 1,828
earlier model28.528446%530 / 1,828513 / 1,828

Stored reports · two-seed means. Paired language views. Not a game win rate.

barga correct: 873 / 833 of 1828; earlier model correct: 530 / 513 of 1828. Training seeds 17 / 18. 914 paired test cases.

What to keep in mind. Card-play question accuracy regressed and remained below the option-position baseline. The implemented variant is must-trump v2.

निर्णयको मापन

Typed decisions

2,000 held-out questions

2,000 held-out English questions per seed.

Typed decisions: held-out question accuracy, two training seeds
CheckpointTwo-seed meanSeed 17: correctSeed 18: correct
barga76.050000%1,527 / 2,0001,515 / 2,000
earlier model76.025000%1,518 / 2,0001,523 / 2,000

Stored reports · two-seed means. Training split used; test excluded.

barga seed 17: 76.35%; seed 18: 75.75%. earlier model seed 17: 75.90%; seed 18: 76.15%. The mean difference is 0.025 percentage points.

What to keep in mind. This is task-specific test accuracy, not zero-shot or broad reasoning performance.

नेपालीमा संवाद

Nepali call decisions

Nepali transcripts · English questions

2,589 held-out questions per seed across 94 held-out calls. Transcript states include Nepali and Nepali–English mixtures; the questions are in English.

Nepali call decisions: held-out question accuracy, two training seeds
CheckpointTwo-seed meanSeed 17: correctSeed 18: correct
barga76.554654%1,993 / 2,5891,971 / 2,589
earlier model76.979529%2,006 / 2,5891,980 / 2,589

Stored reports · two-seed means. Teacher labels, not human gold.

barga mean: 76.554654%; earlier model mean: 76.979529%. The newer checkpoint is lower on this dataset.

What to keep in mind. Measures fit to this call dataset. No real-call text is included. A native-reviewed Nepali benchmark remains missing.

मापनको सन्दर्भ

Read the fine print

Every result has a context.

These are stored question-level evaluations, not independently rerun matches.

A stronger overall score can hide weaker decision families. Legal actions come from separate rule engines.

barga / earlier model · two training seeds · snapshot 64133e0

The room’s saved synthetic call example is barga, training seed 17, recorded on CPU with four threads on 2026-10-05. It is saved output, not live inference.

Read the recorded barga example metadata (JSON)

ampixa.com · barga model card · pip install barga · Python package source · Apache-2.0