aifeifei798/DualBigLittle-MoE-Qwen3-0.6b

🤗 Hugging Face 来源text-generationapache-2.0891M 参数1.8 GBsafetensors✓ 2 个校验和今天更新
需要做种者 →

DualBigLittle-MoE: A Tri-Tier Asymmetric Architecture Decoupling Multitask Cognitive Interference via Dual VRAM Dense Cores & Streaming Micro-Expert Clusters


1. Abstract & Paradigmatic Shift

Conventional autoregressive Large Language Models (LLMs) and standard Mixture-of-Experts (MoE) architectures are fundamentally constrained by Multitask Negative Transfer (the "Seesaw Dilemma") and the Physical VRAM Capacity Wall:

  1. Cognitive Interference: Forcing orthogonal domains (e.g., creative open-ended literary syntax vs. deterministic symbolic calculus and code) to share a monolithic parameter manifold leads to gradient collision, where optimization in one domain inexorably dilutes fidelity in the other.
  2. Memory Hierarchy Inefficiency: Classical MoEs deploy homogeneous, monolithic FFN blocks across layers, creating memory footprints that rapidly exceed consumer hardware capacity. Naive host-to-device offloading typically collapses token generation throughput down to sub-interactive speeds due to PCIe bus thrashing.

DualBigLittle-MoE shatters these constraints by introducing an Asymmetric Tri-Tier Heterogeneous Memory & Compute Topology, inspired by the tri-cluster physical paradigms of modern ultra-efficient mobile SoCs:

  • Tier 1 (GPU VRAM — Arts & Linguistic Anchor Core): The baseline pretrained dense MLP is 100% physically frozen in GPU memory, providing a zero-entropy, non-degrading anchor for general human conversation, grammar, and literary fluency.
  • Tier 2 (GPU VRAM — STEM Specialized Twin Core): An in-situ cloned dense MLP specialized through contrastive differential training ($lr = 2\times 10^{-5}$) to handle deterministic algorithmic logic and mathematical derivations with zero PCIe communication latency.
  • Tier 3 (Host Pinned RAM — 896 Micro-Expert Pool): 896 modular, fine-grained micro-experts (Rank-16 LoRA adapters, precisely 64.0 KB each, 56.00 MB total across 28 layers) residing in host DDR5 RAM, streamed dynamically via non-blocking PCIe DMA through persistent double-buffered staging pipelines.

2. Mathematical Formulation & Forward Execution Dynamics

Given an input activation tensor $\mathbf{x} \in \mathbb{R}^{B \times S \times D}$ at layer $l$, the layer output is synthesized through a multi-stage, hierarchical manifold projection:

========================================================================================
                          TRI-TIER HIERARCHICAL TOPOLOGY
========================================================================================

                                [ Input Activation: x ]
                                           │
               ┌───────────────────────────┴───────────────────────────┐
               ▼                                                       ▼
 ┌───────────────────────────┐                           ┌───────────────────────────┐
 │ MACRO ROUTING CONTROLLER  │                           │ MICRO ROUTING CONTROLLER  │
 │ W_macro: R^{D -> 2}       │                           │ W_micro: R^{D -> 32}      │
 └─────────────┬─────────────┘                           └─────────────┬─────────────┘
               │                                                       │
               ▼ Softmax Gating                                        ▼ Top-8 Selection
        [ w_arts, w_sci ]                                   { i_1, ..., i_8 } in {0..31}
               │                                                       │
               ├──────────────────────────┐                            │ Asynchronous DMA
               ▼                          ▼                            ▼ (Ping-Pong Buffer)
 ┌──────────────────────────┐ ┌──────────────────────────┐ ┌──────────────────────────┐
 │ TIER 1: Arts Anchor Core │ │ TIER 2: STEM Twin Core   │ │ TIER 3: Host RAM Pool    │
 │ Pretrained Native MLP    │ │ Specialized Cloned MLP   │ │ 32 Micro-Experts / Layer │
 │ (100% Frozen in VRAM)    │ │ (VRAM Resident, 0 PCIe)  │ │ (Rank-16, 64 KB each)    │
 └─────────────┬────────────┘ └───────────┬──────────────┘ └───────────┬──────────────┘
               │                          │                            │
               └────────────┬─────────────┘                            │ Top-8 Weighted
                            ▼                                          ▼ Vector Sum
                  [ Dense Base: Y_big ]                       [ Sparse Delta: Y_micro ]
                            │                                          │
                            └────────────────────┬─────────────────────┘
                                                 ▼
                             [ Synthesized Output: Y_big + γ * Y_micro ]
========================================================================================

Stage I: Macro-Cognitive Routing & Dense Dual-Core Fusion

The Macro-Router projects the hidden representation to compute soft allocation probabilities across the physical VRAM cores:

$$ \begin{aligned} \mathbf{z}{\text{macro}} &= \mathbf{W}{\text{macro}} \cdot \mathbf{x} \quad \in \mathbb{R}^{B \times S \times 2} \ [w_{\text{arts}}, w_{\text{sci}}] &= \text{Softmax}(\mathbf{z}_{\text{macro}}, \text{dim}=-1) \end{aligned} $$

The dense backbone computation executes concurrently inside GPU VRAM at native Tensor Core throughput with zero interconnect overhead:

$$ \mathbf{y}{\text{big}} = w{\text{arts}} \cdot \text{FFN}^{\text{Arts}}(\mathbf{x}) + w_{\text{sci}} \cdot \text{FFN}^{\text{STEM}}(\mathbf{x}) $$

Stage II: Micro-Expert DMA Streaming & Low-Rank Projection

Concurrently, the Micro-Router dispatches the Top-$k$ ($k=8$) specialized adapters from the 32-expert pool:

$$ \begin{aligned} \mathbf{z}{\text{micro}} &= \mathbf{W}{\text{micro}} \cdot \mathbf{x} \quad \in \mathbb{R}^{B \times S \times 32} \ \mathcal{K} &= \text{Top-}k(\mathbf{z}{\text{micro}}, k=8) \ \omega_j &= \text{Softmax}(\mathbf{z}{\text{micro}}[\mathcal{K}])_j \quad \forall j \in \mathcal{K} \end{aligned} $$

Each Tier-3 micro-expert computes a low-rank residual transformation via its decoupled bottleneck matrices:

$$ \text{Expert}_j^{\text{micro}}(\mathbf{x}) = \left( \mathbf{W}_B^{(j)} \cdot \left( \mathbf{W}_A^{(j)} \cdot \mathbf{x} \right) \right) \cdot \frac{\alpha}{r} $$

Where:

  • $\mathbf{W}_A^{(j)} \in \mathbb{R}^{r \times D}$, $\mathbf{W}_B^{(j)} \in \mathbb{R}^{D \times r}$, with $D = 1024$ and intrinsic rank $r = 16$.
  • Precision format: Native bfloat16 (2 bytes per parameter).

$$ \text{Parameter Footprint} = (16 \times 1024 + 1024 \times 16) \times 2\text{ bytes} = 65,536\text{ bytes} = \mathbf{64.0\text{ KB}} $$

Stage III: Final Composite Synthesis

$$ \mathbf{y}{\text{final}} = \mathbf{y}{\text{big}} + \gamma \cdot \sum_{j \in \mathcal{K}} \omega_j \cdot \text{Expert}_j^{\text{micro}}(\mathbf{x}) $$

Where $\gamma = 0.3$ represents the empirical residual modulation scalar, preserving base stability while injecting modular domain capacity.


3. Systems Engineering: Zero-Jitter CUDA Double-Buffering

Streaming 224 micro-experts per token ($28\text{ layers} \times 8\text{ Top-K}$) across PCIe without degradation requires deep systems-level hardware co-design:

Host Pinned RAM (DDR5)                      GPU VRAM Staging
┌────────────────────────┐      PCIe DMA    ┌──────────────────────────────┐
│ Paged-locked 56 MB     │  ─────────────►  │ Buffer 0 (Active Compute)    │
│ Micro-Expert Tensor    │                  ├──────────────────────────────┤
│ Storage                │  ─────────────►  │ Buffer 1 (In-Flight Stream)  │
└────────────────────────┘                  └──────────────────────────────┘
                                                 ▲ Ping-Pong Synchronization
                                                 └─ cudaStreamWaitEvent()
  1. Static Memory Allocation (No Allocator Churn): Rather than dynamically invoking .to(device) (which triggers PyTorch CachingAllocator thrashing and memory pool fragmentation), DualBigLittle-MoE instantiates persistent static staging buffers (num_staging_buffers: 2).
  2. Ping-Pong Buffer Overlapping: While Buffer 0 computes GEMM operations on the default compute stream, Buffer 1 asynchronously ingests subsequent expert payloads on self.transfer_stream via non-blocking DMA.
  3. Deterministic Numerical Stability: Measured across 40 identical forward passes, numerical cross-stream jitter was reduced from $1.56 \times 10^{-2}$ to identically $0.0000$, with an end-to-end DMA transfer latency of 0.214 ms (p50: 0.209 ms, p95: 0.238 ms).

4. Empirical Evaluation: The Pareto-Frontier Breakthrough

To demonstrate that DualBigLittle-MoE transcends the multi-task negative transfer barrier, we evaluated the model against the pristine foundation baseline (Qwen/Qwen3-0.6B) across domain-specific validation suites:

Comprehensive Cross-Domain Perplexity (PPL / Eval Loss)

Domain Test Suite Pristine Qwen3-0.6B Baseline DualBigLittle-MoE (Ours) Relative Optimization Metric
Code & Algorithms $7.58$ $6.80$ $-10.4%$ 🚀
Math & Symbolic Logic $6.27$ $5.63$ $-10.3%$ 🚀
Arts & General Prose $8.90$ $8.20$ $-7.8%$ 🚀

Empirical Implication: Unlike classical fine-tuning—where STEM performance gains cause catastrophic degradation in humanistic prose—DualBigLittle-MoE establishes a strict Pareto improvement. All three domains improved simultaneously by $8%\sim10%$, proving that physical dual-core decoupling completely insulates the model from multitask gradient interference.

Macro-Gating Neural Bifurcation Telemetry

Task: Python Quicksort Implementation
══════════════════════════════════════════════════════════════════════
🧠【Dual-Core Energy Allocation (Tier-1 & Tier-2 GPU Resident)】:
   🏛️  Tier-1 Arts Anchor Core:   2.0% [                    ]
   🔬 Tier-2 STEM Twin Core:     98.0% [███████████████████ ] <--- 98.0% Autonomous Handover
──────────────────────────────────────────────────────────────────────
🧩【Micro-Expert Allocation (Tier-3 Streaming)】:
   💻 Code Specialists:          29.6% (8,095 calls)
   🧮 Math Specialists:          25.5% (6,957 calls)
   ✍️  Writing Specialists:       44.9% (12,276 calls)
══════════════════════════════════════════════════════════════════════

Task: Jiangnan Rainy Alley Creative Prose
══════════════════════════════════════════════════════════════════════
🧠【Dual-Core Energy Allocation (Tier-1 & Tier-2 GPU Resident)】:
   🏛️  Tier-1 Arts Anchor Core:  78.8% [███████████████     ] <--- Re-established Control
   🔬 Tier-2 STEM Twin Core:     21.2% [████                ]
──────────────────────────────────────────────────────────────────────
🧩【Micro-Expert Allocation (Tier-3 Streaming)】:
   💻 Code Specialists:          28.9% (9,466 calls)
   🧮 Math Specialists:          24.7% (8,080 calls)
   ✍️  Writing Specialists:       46.3% (15,158 calls)
══════════════════════════════════════════════════════════════════════

5. Architectural Specification & Hyperparameter Configuration

{
  "model_type": "dualbig_moe",
  "architectures": ["DualBigMoEForCausalLM"],
  "auto_map": {
    "AutoConfig": "configuration_dualbig_moe.DualBigMoEConfig",
    "AutoModelForCausalLM": "modeling_dualbig_moe.DualBigMoEForCausalLM"
  },
  "base_model_name_or_path": "Qwen/Qwen3-0.6B",
  "expert_pool_location": "host",
  "num_staging_buffers": 2,
  "gamma": 0.3,
  "group_sizes": [8, 8, 16],
  "num_experts": 32,
  "top_k_experts": 8,
  "lora_rank": 16,
  "lora_alpha": 16.0,
  "num_hidden_layers": 28,
  "hidden_size": 1024,
  "intermediate_size": 3072,
  "num_attention_heads": 16,
  "num_key_value_heads": 8,
  "vocab_size": 151936,
  "max_position_embeddings": 40960
}

6. Quickstart: Native Hugging Face Integration

DualBigLittle-MoE is fully packaged for the Hugging Face Hub. Any environment with standard transformers can ingest and instantiate the model directly via trust_remote_code=True.

Standard Inference Pipeline

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "aifeifei798/DualBigLittle-MoE-Qwen3-0.6b"

# Load tokenizer and model via native dynamic code execution
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="cuda:0"
).eval()

# Construct standard ChatML prompt
messages = [
    {"role": "system", "content": "You are a precise, logically rigorous assistant."},
    {"role": "user", "content": "Write an optimized Python quicksort function with boundary safety."}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to("cuda:0")

# Reset telemetry buffers prior to execution
model.reset_routing_stats()

# Execute generation
with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=256,
        temperature=0.7,
        top_p=0.9,
        repetition_penalty=1.15
    )

# Decode output and inspect real-time routing report
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(response)

# Output detailed multi-tier allocation metrics
model.routing_report()

Hardware Deployment Modes

In config.json, the execution topology can be toggled via a single flag:

  • "expert_pool_location": "host" (Default): Activates the 56 MB DDR5 streaming pipeline. Peak VRAM footprint: ~1.70 GB. Ideal for edge hardware, laptops, and consumer cards.
  • "expert_pool_location": "device": Pins all 896 micro-experts permanently in GPU memory. Eliminates PCIe transfers entirely, maximizing throughput on high-VRAM platforms (RTX 4090 / 5090).

7. Comparative Systems Benchmarking

Metric / Dimension Traditional Dense Base Monolithic MoE (Mixtral Style) DualBigLittle-MoE (Ours)
Active Knowledge MLP Capacity $2.64\times 10^8$ params $N \times \text{FFN}$ (Massive) $5.57\times 10^8$ params (+111.7%)
VRAM Footprint (FP16/BF16) $1.14\text{ GB}$ $> 14.5\text{ GB}$ $1.70\text{ GB}$
Host System RAM Overhead $0.00\text{ MB}$ $0.00\text{ MB}$ $56.00\text{ MB}$
Multitask Gradient Decoupling None (Seesaw trade-off) Partial / Entangled Physical Isolation (98% Bifurcation)
PCIe Stream Transfer Latency N/A High (Bus saturation) $0.214\text{ ms}$ (Ping-Pong DMA)
Sustained Decoding Speed $\sim 22.0\text{ tok/s}$ Constrained on Consumer VRAM $26.1\text{ tok/s}$

8. Citation & Prior Art

If you incorporate the DualBigLittle-MoE architectural topology, macro-micro hierarchical routing, or asynchronous micro-expert streaming pipelines into your research, please cite:

@misc{dualbiglittle_moe_2026,
  author = {aifeifei798 and Community Contributors},
  title = {{DualBigLittle-MoE: A Tri-Tier Asymmetric Architecture Decoupling Multitask Cognitive Interference via Dual VRAM Dense Cores and Streaming Micro-Expert Clusters}},
  year = {2026},
  publisher = {GitHub and Hugging Face},
  howpublished = {\url{https://github.com/aifeifei798/DualBigLittle-MoE}},
  note = {Hugging Face Checkpoint: \url{https://huggingface.co/aifeifei798/DualBigLittle-MoE-Qwen3-0.6b}, Transformers Issue \#49183}
}