ruvnet/ruos-foundry-tools-qwen3-30b-a3b-e64

🤗 Hugging Face sourcetext-generationapache-2.016B params32 GBsafetensorsHF checksums availableupdated today
No torrent yet

ruos-foundry-tools-qwen3-30b-a3b-e64

ruos-foundry-tools-qwen3-30b-a3b-e64 is a ruOS specialist cut from Qwen/Qwen3-30B-A3B-Instruct-2507 with MoE-Foundry (ADR-064): the parent's routed experts were traced on the ruOS ruos-tools calibration set, the 64 most-used experts per routed layer were kept, and a complete smaller Mixture-of-Experts model was exported — same backbone, same tokenizer, same token top-k, fewer experts per layer. No weight was changed; experts were removed and renumbered and the router rows sliced to match.

Status: unevaluated and disabled

Per MoE-Foundry's rule an exported specialist starts quality_status: unevaluated, enabled: false. A structural export proves tensor integrity, not retained capability. This checkpoint is published so that slim-eval (slim/eval: run, redteam, regression, verdict) can measure it against the parent on the frozen ruOS test split; until that verdict is recorded here it must not be routed to. The full parent remains the fallback.

Parent

repo Qwen/Qwen3-30B-A3B-Instruct-2507
revision 0d7cf23991f47feeb3a57ecb4c9cee8ea4a17bfe
licence Apache-2.0
MoE-Foundry parent_id 44c2b6971fe6979409ac72586db5fa4abfd73746afc52607368a59740848149d
experts per routed layer (parent → this) 128 → 64
routed layers 48
token top-k (unchanged) 8

Calibration

Domain ruos-tools: 200 rows (200 from the val split, 0 topped up from train; the test split is never used for calibration), ~61028 tokens, families {"stack_qa":75,"tool_routing":125}, licences {"MIT":188,"project-owned":12}. Calibration file sha256 94a5527b141269ab5078a14e3d129b387aa69412c65334160715fd9a58dbfef5. Texts are the ChatML prompt plus the reference answer.

Router traces: 412 tasks, 148144 tokens, 7110912 rows on NVIDIA A100-SXM4-80GB (bfloat16, transformers 4.51.3); trace sha256 7aabb05cbf715156437393fc5a5b0d756951ecee35eb4542a06fc812f11f9403.

Selection method: mass (accumulated routing probability per expert per layer) — a usage proxy, not causal importance.

Receipt

specialist_id 08bfb95ad23a8aa9374fc70519ab57f65271db938196f4dc160b58ff50e284b4
checkpoint_id 6207a9d03a49b4329c6daa176846360e6e63d9333a6b98e54c5eda9b60f360ee
mask_sha256 8140132e55c5e22c8f7757a0df08260e49f5030e149ee24c7f17a108d367e4fe
parent_id 44c2b6971fe6979409ac72586db5fa4abfd73746afc52607368a59740848149d
input tensor bytes 61064245248
output tensor bytes 32060633088
reduction 47.5%

separator_receipt.json in this repo is the full MoE-Foundry receipt including the per-layer expert mask.

Files

path bytes sha256
LICENSE 11343 05cab46843576551502bfdf712f84e93e6e9590d9997306ed4f6635ef82811d9
SHA256SUMS 737 92aecff9bddf335e246ee0346d6b9b19e1631a2a6f3fe13e7285e36f248bf452
config.json 963 f0bf60e7da89b14b5763344441e4f60eccda09f62c0111c5f178283c5f3d81ec
generation_config.json 239 19d306dd769db12a9d710b44cf7f83b635efbe5166b84fb4358a08fb7d88bb53
merges.txt 1671839 599bab54075088774b1733fde865d5bd747cbcc7a547c5bc12610e874e26f5e3
model.safetensors 32061839304 3604df06bbf46834ffe89b14a8570173fcaf4096dc0e2cdf7e3e81bf557e513c
separator_receipt.json 32984 2aa23b33133b58e5d7c071c4bf62622366f3f655298715e055f5619ca4a89314
tokenizer.json 11422654 aeb13307a71acd8fe81861d94ad54ab689df773318809eed3cbe794b4492dae4
tokenizer_config.json 9377 a62ff0a2472a0fa1b8eaabcb57c59b58afa42a22831dc141400b6e0cf2b65ce3
vocab.json 2776833 ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910

Run

vllm serve ruvnet/ruos-foundry-tools-qwen3-30b-a3b-e64 --max-model-len 8192

Loads with transformers as a standard qwen3_moe checkpoint (single safetensors file, num_experts reduced in config.json).

Limitations

  • Unevaluated: no capability, memory or latency claim is made here.
  • Retained experts were chosen by routing mass on ruOS calibration prompts; requests outside that domain should go to the parent.
  • Memory: fewer experts means a smaller checkpoint; loading several specialists next to the parent can use more total memory than the parent alone.

Provenance

  • MoE-Foundry 6677a25 (moe-separator inspect → profile_hf → select → export → mixture)
  • run foundry-20260907T172935Z-qwen3-30b-a3b on a single vast.ai GPU; ruos-desktop slim/foundry + slim/scripts/foundry-e2e.sh
  • authorisation: rUv, "implement this using ruvnet/MoE-Foundry using vast.ai in a worktree, implement e2e and push models to repo"