FINAL-Bench/Darwin-Chimera-4B-Gen1

🤗 Hugging Face 来源text-generationapache-2.04B 参数8.0 GBsafetensors✓ 2 个校验和今天更新
一条命令提交

在你的模型文件夹旁边运行它。它会制作种子、将文件与 Hugging Face 比对,然后提交。你只需开始做种,并粘贴你账户中的密钥。它只读取你的文件,绝不修改。如果愿意,可以先阅读脚本。

curl -fsSL https://pirateface.co/package.sh | bash -s -- --repo FINAL-Bench/Darwin-Chimera-4B-Gen1 ./model-folder
需要做种者 →

Darwin-Chimera-4B-Gen1 (Backbone · Research)

⚠️ Generation-1 backbone — a research checkpoint, not a product. Private repo. This is a Qwen3-4B derivative, not a from-scratch model. We state this explicitly.

What this is

Darwin-Chimera-4B-Gen1 is the first-generation adapter backbone of the Darwin-Chimera line. We take Qwen/Qwen3-4B and re-wire only its attention via VIDRAFT attention-healing, while freezing the FFN, embeddings, and lm_head, and convert the attention to a sliding-window configuration.

The purpose is to verify that a VIDRAFT-healed attention circuit can sit on a frozen knowledge core — the foundation for Generation-2 (FFN cross-breeding with other models).

Honest weight fingerprint (vs Qwen/Qwen3-4B)

Measured relative change ||A−B|| / ||A|| against the original Qwen3-4B:

Component Relative change Note
FFN (mlp) 0.000% frozen — identical to Qwen3-4B
embed / lm_head 0.000% frozen — identical
attention (self_attn) 3.0% mean (7.5% max) healed
layernorm 0.04% minimal
config (hidden/inter/layers/vocab) identical only sliding_window=4096 added

→ At the weight level this checkpoint is clearly a Qwen3-4B derivative. We make no claim of independence or from-scratch training. Knowledge/FFN is 100% Qwen3-4B.

Training

  • Method: attention-only healing (self_attn + per-layer norms trainable; FFN/embed/lm_head frozen)
  • Attention: full → sliding window (4096), 5:1 sliding:full layer ratio
  • Tokens: ~3B (Korean-centric annealing mix)
  • Base: Qwen/Qwen3-4B (Apache 2.0)

Evaluation (base, zero-shot — reference only)

  • Generation: 6/6 domains coherent (Korean / English / science / code / math / biology), no gibberish
  • KMMLU (6 subjects, 240Q, zero-shot, greedy): 27.1% vs Qwen3-4B base 13.3% (same protocol, +13.8pp)
  • Absolute KMMLU is low because this is a base (non-instruct) checkpoint; instruction-following and benchmark quality are expected to come from a later SFT stage. The comparison above is a same-condition relative measurement, not an absolute SOTA claim.

Intended use

  • Backbone for Darwin-Chimera Generation-2 (cross-architecture FFN cross-breeding research)
  • Research and experimentation only. Not for production.

License & attribution

Apache 2.0, inherited from Qwen/Qwen3-4B. Built on Qwen/Qwen3-4B.