Joakimpalm-Zen/Xyntetik-Blad-Genesis-0.7M

🤗 Hugging Face sourceapache-2.025 MBother✓ 9 checksumsupdated today
Needs seeder →

Xyntetik-Blad-Genesis-0.7M

Research artifact, not a general-purpose release.

  • What it is: three frozen 686,756-parameter program-proposal cores (seeds 29093641, 29093642 and 29093643, each at 160, 320 and 640 meta-iterations) from GENESIS generation I, study GL-003c. Each reads 5 input/output examples in one synthetic 35-opcode list DSL and proposes programs in that DSL; the host executes them and keeps the first that fits all 5.
  • Why it exists: these are the cores behind the record's measurements (GL-003c, GL-003d, GL-003e's grammar-1 comparison, GL-003f's small arm, GL-005's starting point and GL-007's base), published so that every number can be checked against the weights.
  • Parent model: none. Trained from scratch; no human text in the data, the labels, the reward or the reader. With no parent there is no KLD fidelity bar, and none is claimed. This is not an Xyntetik Runner model.
  • Useful for normal inference? No. It does not read or write text. It proposes programs in its own toy DSL, and about half of its spec-consistent answers fail on further inputs.
  • External replication: failed (GII-001); see the training record.
  • Experiment: collection Xyntetik Research: GENESIS (https://huggingface.co/collections/Joakimpalm-Zen/xyntetik-research-genesis-6abcfa50b6707dec9bfee356)
  • Evidence: training record (dataset: the full Genesis I record, with 18 preregistrations, their amendments, decisions, reports, raw archives, code, the synthesis and the continuous-medium checkpoints). Not copied here; the three GL-003c checkpoint sets are also inside its runs/GL-003c/GL-003c-raw-ckpt-*.tar.gz.
  • Read first: GENESIS I: what survived 18 preregistered experiments on machine-only learning, the article, kept in the training record.

The cost first. At equal CPU-seconds, with its training charged, this core was level with the best blind search it was compared with (a pure-Python genetic algorithm) over 450-task streams, and behind at 160 iterations. It led in one confirmatory study only, by +0.037 [+0.018, +0.056] over a single 1,800-task stream; that lead was not resolved on the first 1,200 tasks or at size 8, and the genetic algorithm overtakes it above 85 CPU-seconds per task. About half of the programs it returns as matching all 5 examples fail the 20 hidden inputs. A 4× larger core lost to more iterations of this one at equal FLOPs. Online updates while running (GL-005) and a self-extracted vocabulary (GL-007) were both killed. What held up: at equal executions (50,000 per task) it solves more held-out size-6 to 8 tasks than that genetic algorithm, in every study that measured it after the first attempt was killed.

Quickstart

CPU-only PyTorch is enough. Nothing else is needed.

pip install torch
hf download Joakimpalm-Zen/Xyntetik-Blad-Genesis-0.7M --local-dir Xyntetik-Blad-Genesis-0.7M
cd Xyntetik-Blad-Genesis-0.7M
sha256sum -c SHA256SUMS
python run_example.py

Tested on this release's staging box on 2026-09-30 with taskset -c 0-15 /home/lab/vllm-venv/bin/python run_example.py (AMD Ryzen Threadripper 9980X, CPUs 0-15, torch 2.11.0+cu130 with CUDA hidden, 4 threads). Output:

weights  weights/seed-29093641/model_it640.pt  sha256 66e10e3f62717cf7…  686,756 parameters  finite=True
target   C3 X FILT>0 SORT REV TAKE SUM   (hidden from the core; only the 5 examples below are shown to it)
  in [8, 10, -14, 0, 16, 15, 9, 3]      out 41
  in [6, -3, 16, -8, 2]                 out 24
  in [-10, 0, -7]                       out 0
  in [-10, -12, 5, 14]                  out 19
  in [-10, 6, 11, 4, -3, 14]            out 31
search   SBS k=1000: 1000 distinct programs sampled, 345 executed (1725 executions), stop=match, 0.34 CPU-s, 0.09 s wall
result   C3 X SORT REV TAKE FILT>0 SUM   (draw 345, consistent with all 5 examples)
check    agrees with the target on 1000 of 1000 fresh random inputs (sampled agreement, not proven equivalence)

The target (the sum of the three largest positive values) is an illustrative size-7 task chosen for this card, not a measured one. The core found a different program with the same behaviour on the 1,000 inputs checked. The same task with three other seeds (--seed 1, 2, 3; each seed draws new example inputs and a new search stream) shows the limits:

seed 1: result   X FILT>0 SUM   (draw 14, consistent with all 5 examples)
        check    agrees with the target on 764 of 1000 fresh random inputs
seed 2: result   no spec-consistent program in this sample (1,000 programs, 5,000 executions)
seed 3: result   C3 X REV FILT>0 TAKE MAP*2 MAP/2 SUM   (draw 156, consistent with all 5 examples)
        check    agrees with the target on 799 of 1000 fresh random inputs

Two of the three answers fit the 5 examples and are wrong, and the third search finds nothing within its budget. That is the record's finding in miniature: "solved" means agreement on sampled behaviour, and about half of the claims fail on further inputs.

Run it

python run_example.py --target "C3 X SORT REV TAKE SUM"      # build the 5 examples from any DSL program
python run_example.py --spec my_task.json                    # {"examples": [{"input": [3, -1, 7], "output": 9}, ...]}
python run_example.py --weights weights/seed-29093642/model_it160.pt --k 10000 --seed 5
  • --k is the SBS sample size. The studies used k = 10,000, the number of distinct programs a 50,000-execution budget admits on 5 examples. The default here is k = 1,000.
  • A spec needs 5 examples. Inputs are lists of 2 to 8 integers in [−16, 16]. Outputs are an integer or a list, and integer results saturate at [−256, 255]. Other inputs are outside everything the core saw.
  • Programs are postfix: C3 X SORT REV TAKE SUM pushes 3, pushes the input list, sorts it, reverses it, takes 3, and sums. The 35 opcodes and their semantics are in code/gl003_dsl.py (OPCODES).
  • With --target, the script also runs the found program and the target on --fresh random inputs (default 1,000). With --spec, only consistency with the 5 examples is known.

What it is, exactly

The domain. Typed postfix bytecode over two value types, an integer and a list of integers, run by a stack VM (code/gl003_dsl.py). It has 35 opcodes (constants 1-3, the input list, reverse, sort, element-wise maps, filters, prefix scans, head/last/min/max/sum/length/counts, take/drop/index, element-wise zips, add and subtract). Programs have at most 10 opcodes. Tasks are the input-output behaviours of random programs drawn from a Dirichlet(2)-weighted grammar over those opcodes: synthetic, generated, and with no human data. A search sees 5 examples; the judge scores 20 hidden inputs.

The model (code/gl003_learn.py, Model). A decoder-only pre-LN transformer: d_model 128, 3 layers, 4 heads, MLP 512 with GELU, 128 positions, 686,756 parameters, float32. It has 552 input tokens (PAD, BOS, IN, OUT, OUTI, one token per integer value from −256 to 255, and 35 opcodes) and 36 output tokens (the 35 opcodes and END). The spec is encoded as IN x1..xn OUT y1..ym (or OUTI y) per example, then BOS, then the program. Decoding is masked to type-valid continuations, so every sampled program is well formed.

Training: S1b hindsight relabelling, 640 meta-iterations. In each meta-iteration the model samples 64 programs for each of 64 training tasks (temperature 1, type-masked, length limit 4, 6, 8, then 10 by iteration). The host executes every distinct program on the task's inputs. The outputs the host computed become the specification that program solves: a replay buffer keeps the shortest program per (task, output vector), and 32 AdamW steps of batch 256 maximise the likelihood of buffer programs given their relabelled specifications (lr 1e-3, betas 0.9 and 0.98, weight decay 0.01, 200 warm-up steps, clip 1.0). The training tasks' own outputs are not a training signal, and there is no reward. The record's outcome-reward arm stayed within noise of uniform random (+0.005 [−0.009, +0.021] in GL-003b). Training ran on 9,945 training task specs, on CPU with 2 threads per learner, for about 27,600-27,900 CPU-seconds per seed (7.7 CPU-hours; GENESIS.json has the per-checkpoint figures).

Test-time search (code/gl003b_learn.py, sbs, search_task_sbs). Stochastic Beam Search (Kool, van Hoof and Welling 2019) draws an exact sample of k distinct programs without replacement. They are executed in sampling order through a budgeted executor, and the first program matching all 5 examples is the answer. The first answer is final. This duplicate-free sampler is part of the result: it alone was worth +0.053 [+0.021, +0.085] in GL-003b's ablation.

Files.

  • weights/seed-<s>/model_it{160,320,640}.pt: PyTorch state dicts, 2.76 MB each. They load with torch.load(..., weights_only=True).
  • SHA256SUMS: the nine files, identical to the record's runs/GL-003c/checkpoint_sha256.txt.
  • code/: the three harness modules, unchanged from the record.
  • run_example.py: this card's runner.
  • GENESIS.json: architecture, recipe, checkpoint hashes, source commits and the headline measurements with their files.

Why

GENESIS asked whether a machine could learn a medium of its own with no human anywhere: not in the data, the labels, the reward or the reader. The study's thesis is a native AI language made for AI. The track made execution truth the judge: what the host computes when it runs a program. Its first stage, S1, asked whether a tiny model trained from scratch with no human text can learn to produce programs that do what a task asks, beating the best blind search at equal executions in a band where that search solves under half.

S1 (GL-003) was killed by its first kill criterion. S1b (GL-003b), with a duplicate-free sampler and a longer horizon, passed on fresh seeds and tasks. GL-003c trained S1b's recipe to 640 meta-iterations on fresh seeds, and those are the cores published here. Later studies used them as frozen cores (GL-003d, GL-003e, GL-003f), as starting points (GL-005) and as a base (GL-007). Of the track's kill-switch stages (S1, S1b, S2, S2b, S3) plus GL-007, one passed and five were killed.

The owner froze generation I on 2026-09-30 and chose to publish it as this model and an evidence dataset. The cores are a toy-scale research artifact, the one component of the programme that held up, released so that the claims and the case against them can be checked. The record's own prior-art reading calls the result a tiny-scale DeepCoder / CrossBeam / CodeIt replication, and no novelty is claimed.

Measured envelope

All figures come from the preregistered studies. Intervals are each study's registered 95% intervals. The file paths are in the training record under genesis-language/. The comparator throughout is GA-64-graded, a pure-Python genetic algorithm, the best of 11 blind variants in four families as chosen on calibration tasks. The task band is sizes 6-8 (GL-005: 5-7).

At equal CPU-seconds, training charged (the cost)

Study Weights Reading Verdict Report
GL-003c Q2 (exploratory) these, 160 / 320 / 640 Training charged over 450 tasks: −0.045 [−0.088, −0.004] at 160 (blind search ahead); +0.014 [−0.024, +0.053] at 320 and +0.019 [−0.021, +0.058] at 640 (level) behind, then level report/GL-003c.md
GL-003d (confirmatory) these, 640, frozen Whole training CPU charged to one 1,800-task stream: 0.438 against 0.401 at 30.75 CPU-s per task, E = +0.037 [+0.018, +0.056] (pre-data projection +0.065). At n = 1,200 tasks +0.021 [−0.003, +0.046]; the lower bound clears zero first at n = 1,500. Size 8 alone: +0.022 [−0.012, +0.056]. GA's solve rate passes the core's from 85.0 CPU-s per task "core ahead at equal CPU-seconds", confirmed, narrowly report/GL-003d.md

Scale and adaptation (both against)

Study Weights Reading Verdict Report
GL-003f (characterization) these as the small arm; a 4× larger core trained for the study At equal training FLOPs: −0.063 [−0.101, −0.024] at the full budget (0.381 against 0.444); −0.146 at ¼ and −0.112 at ½ "size hurts" report/GL-003f.md
GL-005, S3 reframed (confirmatory) these at 160, as starting cores Online hindsight updates from verified successes: 0.520 against the frozen core's 0.536 (−0.016 [−0.041, +0.009]). K1 −0.030 [−0.057, −0.006] on measured CPU, −0.045 [−0.073, −0.020] load-robust. K2: two late promotions doubled regression failures, +0.079 [+0.031, +0.126]. Memory − frozen: +0.008 [−0.011, +0.028] killed by K1 and K2; the memory rule fired report/GL-005.md
GL-007 (confirmatory) these at 640, trained 160 further iterations with and without extracted words Library 0.448 against plain 0.457, D1 −0.009 [−0.026, +0.007]; all 48 words are input-plus-one-operation bigrams killed by K1 (K2 also fires) report/GL-007.md

At equal executions (50,000 per task)

Study Weights Reading Verdict Report
GII-001, Genesis II (confirmatory; external, human-written tasks) these at 640, plus 3 fresh seeds of the same recipe 127 human-written list functions (LambdaBeam's handwritten tasks, Rule et al.'s list functions), "solved" = agreement with a reference on up to 77,437 inputs: 0.247 against GA-512-graded's 0.249, the best of 12 blind searches, −0.002 [−0.039, +0.036]. At equal CPU with training charged: −0.047 [−0.080, −0.015] (behind). In-distribution control: +0.273 [+0.232, +0.314] (reproduces). Secondaries: LambdaBeam +0.073 [+0.020, +0.133]; Rule et al. −0.053 [−0.101, −0.010] killed report/GII-001.md
GL-003b, S1b (confirmatory) sibling seeds 29092641-43 of the same recipe at 160 (not in this repo) 0.393 against 0.211: K1 +0.182 [+0.137, +0.228]; K2 +0.384 [+0.341, +0.427] over uniform random. Ablations: sampler +0.053 [+0.021, +0.085], horizon +0.079 [+0.046, +0.113] passed report/GL-003b.md
GL-003c Q1 (exploratory) these, 160 / 320 / 640 450 fresh band tasks: 0.336 [0.295, 0.376] / 0.416 [0.375, 0.457] / 0.444 [0.402, 0.485], against GA 0.187 [0.157, 0.219]; leads +0.148 / +0.228 / +0.256. D(640) − D(160) = +0.108 [+0.074, +0.144]; 320 → 640 added only +0.028 [+0.001, +0.056] "grows", with diminishing returns report/GL-003c.md
GL-003d (descriptive) these, 640 +0.276 [+0.254, +0.298] over GA on the 1,800-task stream not a registered reading report/GL-003d.md
GL-003e (confirmatory) grammar-2 cores by the same recipe; these at 320 as the grammar-1 comparison On a rank-reversed second grammar: 0.451 against 0.199, K1 +0.252 [+0.202, +0.302]. These grammar-1 cores already solve 0.413 there (+0.214 [+0.162, +0.267] over GA), so the grammar-2 core's own edge is +0.038 [−0.001, +0.077] "replicates" report/GL-003e.md

Hidden inputs. Of these cores' spec-consistent claims at 640, 48% fail the 20 hidden inputs in GL-003c and 47% in GL-003d (GA: 59%). Every arm is scored the same way.

Limits, read before quoting

External replication failed (GII-001, 2026-09-30): on human-written list tasks these cores were level with the best blind search at equal executions and behind at equal CPU. On 127 functions from LambdaBeam's handwritten tasks and Rule et al.'s list functions, judged on up to 77,437 inputs each, six cores (these three at 640 and three fresh seeds of the same recipe) solved 0.247 against GA-512-graded's 0.249: −0.002 [−0.039, +0.036]. At equal CPU, with training charged, they were behind: −0.047 [−0.080, −0.015]. In the same run the in-distribution control reproduced (+0.273 [+0.232, +0.314]), so the leads in the Measured envelope belong to the synthetic generator's tasks. As registered secondaries, the cores led on LambdaBeam's 49 tasks (+0.073 [+0.020, +0.133]) and trailed on Rule et al.'s 78 (−0.053 [−0.101, −0.010]). The study cannot rule out a small advantage of about +0.03. There is no rescue study, by design. Report: genesis-language/report/GII-001.md in the training record.

The case against, in full, from the Genesis I synthesis (SYNTHESIS-GENESIS-I.md, section 4, checked point by point against the record). Points 5, 12 and 13 concern the separate medium line (GL-002), whose checkpoints are in the dataset, not here.

  1. A small synthetic DSL and a synthetic lookup task. The core line uses a 35-opcode postfix DSL (at most 10 opcodes, inputs of 2-8 values in [−16, 16], saturation at [−256, 255]), and its tasks are behaviours of random programs from the same grammar family. No human-text knowledge, no pretrained model and no non-synthetic task appears anywhere after GL-001. The prior art warned in advance that a domain with an exact verifier is usually one a program already solves: a win here is evidence about the medium, not a product.
  2. The second grammar was easy for grammar-1 cores: 0.413 against the grammar-2 core's 0.451, +0.038 [−0.001, +0.077]. The data distance was large (operator-frequency JSD 0.172 bits, 1.1% behaviour overlap), but both grammars are Dirichlet(2) draws over the same 35 operators with the same inputs and sizes.
  3. The equal-CPU win is narrow, needs a long stream, and is unresolved at size 8. +0.037 [+0.018, +0.056] against a projected +0.065. Break-even is about 399 tasks, and the lower bound clears zero only at n = 1,500. Over 450 tasks the result was level, or behind at 160 iterations. The claim holds only at the core's own cost level; GA passes it from 85.0 CPU-s per task.
  4. About half of every arm's claims fail the hidden inputs: 45-65% across the GL-003 family, and about a third in GL-005. All solve rates are "first answer final".
  5. Continuous formation failed on 1 of 6 seeds; discrete formation is about 50% (medium line: GL-002d 5 of 6; soft sample 15 of 32, 0.47 [0.31, 0.63]; straight-through 0 of 53).
  6. The larger core hurt: width ×2 (3.97× the non-embedding parameters) with µP learning rates and nothing tuned, −0.063 [−0.101, −0.024] at the full budget and worse at ¼ and ½, with 1.6-1.7× the search CPU. A crossover at a larger budget was not measured.
  7. None of the adaptive mechanisms paid: the per-host cost adapter (1.40× and 1.38× atlas), online weight updates (−0.016 uncharged; K1 fired), the retrieval memory (+0.008 [−0.011, +0.028]) and the machine's own vocabulary (−0.009). The stability rule and its checkpoint half, the warm start, the temperature floor and a wider core did not matter either. Every tested part of the track's "self-evolving per machine" clause was killed. What held up is frozen: a frozen core, searched more.
  8. Many studies were designed after earlier results: 15 of 18. S1b's two changes were named by a post hoc reading of S1's killing data, and GL-003d's long-stream design came from a descriptive curve seen after GL-003c. Each was preregistered before its own data, so within-study inference holds, but the sequence as a whole is adaptive.
  9. Co-tenant load distorted CPU-seconds. S1b's "48× GA's CPU" was mostly OpenMP spin-waiting; GL-003d's core search cost 5.5× its GL-003c price; GL-003e's core work was inflated about 2-3.5×; and in GL-005 the same SBS pass cost about 25% more in one arm's workers than in another's. Mostly the load leaned against the core, but the equal-CPU numbers belong to one night's placement.
  10. CPU-only, one host, toy scale. One Threadripper box shared with other work; 0.69 M parameters. S4 never ran, so portability and per-machine claims are untested.
  11. Prior art already covers parts: DeepCoder, CrossBeam and CodeIt ("a tiny-scale … replication"); SOAR, which self-trains on execution-judged search traces from a pretrained model; FFTW, ATLAS and TenSet for S2; Coconut for the continuous medium, and the theory strand for its constant-factor speed; KernelBench-Verified for the verifier failure; the papers that predicted GL-005's and GL-007's nulls. Codes private to their model (the result behind "adapters may not change the medium") were not tested at all.
  12. The medium is not machine-native in the strong sense: a nearest-class-mean reader decodes it at 0.83-1.00, and what the model chose is a schedule, not an alien code.
  13. The medium line never ran the control the prior art called its only worthwhile form (an invented code against equally RL-trained brevity at matched compute).
  14. The thesis was never tested as a whole. This core still speaks a human-designed DSL; the medium lives on another task and was never used inside the core; portability never ran; both tests of evolution failed.
  15. Human design is everywhere except the data, labels, reward and reader. The DSL, grammar family, curricula, architectures, reading rules and task generators are human choices, so "no human anywhere" is met only in that narrowed sense.
  16. Several passes sit at their minimum margin (GL-002c 2 of 3; GL-002d 5 of 6; GL-002e's control 2 of 3; GL-003e's X5 "not supported").
  17. The blind baseline field was thin. HEAP, expected to be best, solved almost nothing in the band, and the enumerative families scored under 0.01. The core's lead is over one pure-Python GA; nothing stronger (for example A* enumeration) was tried.
  18. Process irregularities are on the record, all disclosed: API session limits stopped the operating agent during ten studies (detached chains finished them); the operator read some data during GL-002c, GL-002f and GL-003d before the relevant amendments; a double launch voided GL-004b's first Phase B, and a build race voided GL-004's first calibration. None changed a registered rule after the data it affects.

And for this model specifically:

  • "Solved" means agreement on sampled behaviour, not proven equivalence: all 5 spec inputs and 20 hidden inputs. On 1,000 fresh inputs, solved programs mostly agree with the target, but not always: means of 0.994-0.996 with minima of 0.779 (core) and 0.849 (GA) in GL-003, minima of 0.814 and 0.807 in GL-003b, and means of 0.990-0.999 later. The equivalence checks are sampled; there is no proof and no exhaustive check over the value range.
  • Not "trained from execution outcomes". The core is trained by hindsight relabelling of its own executed programs, not by a reward on outcomes. The system that leads is the model and the duplicate-free sampler.
  • Three seeds per arm, one DSL, train and test tasks from the same generator family. Genesis I tried no non-synthetic domain; GII-001 then tested human-written external tasks, and the advantage over the best blind search did not survive (see the first limit above).
  • No parent, so no fidelity row. The house KLD bar applies to copies of a parent model; this model has none and claims none. "Parent-agreement is not a capability benchmark" does not apply either way.
  • Not an Xyntetik Runner model. There is no GGUF, no tokenizer and no chat template. It runs only through the PyTorch code here.
  • The Quickstart task is illustrative. Measured rates are the band rates above, at k = 10,000; at the card's default k = 1,000 the search is weaker.

Reproduce

The harness is in the training record under genesis-language/, identical to the shade branch genesis-language at 6bf599e25. GL-003c's full orchestration is gl003c_run.sh, which runs Phase 0 (unit tests), A (tasks and calibration searches, then the band), B (training), C160/C320/C640 (checkpoint searches and judging), G (GA at 50,000 executions and CPU-budgeted) and D (the decision). The steps that make and test these weights:

cd genesis-language
export CUDA_VISIBLE_DEVICES=''            # CPU only; the learner refuses CUDA
PY=python                                 # the study: torch 2.11.0+cu130, Python 3.12
$PY gl003c_tasks.py tasks --out TASKS     # fixed seeds: test 29093633, calibration 29093634, dev 29093635, train 29093636
# Phase A (calibration searches, then gl003c_analyze.py band) writes RUNS/band.json; see gl003c_run.sh
for s in 29093641 29093642 29093643; do
  $PY gl003c_learn.py train --arm hindsight --seed $s --threads 2 --tasks TASKS --out RUNS/learn --max-seconds 32400
done
for t in 160 320 640; do for k in 0 1 2 3; do
  $PY gl003c_learn.py search --run RUNS/learn/hindsight-29093641 --iteration $t --tasks TASKS \
      --band RUNS/band.json --shard $k --nshards 4 --threads 1
done; done                                # likewise for the other two seeds
$PY gl003c_tasks.py judge --tasks TASKS --candidates <search .jsonl> --out <judged .jsonl>
$PY gl003c_analyze.py checkpoint --runs RUNS --tasks TASKS --t 640 > RUNS/ckpt-640.json
$PY gl003c_analyze.py decide --runs RUNS --tasks TASKS > RUNS/decision.json

Check the weights with sha256sum -c SHA256SUMS. The record's hashes are in runs/GL-003c/checkpoint_sha256.txt. Training is seeded per stream, and the study's unit tests check two properties bit-exactly: the weights at iteration t equal those of a t-iteration run with the same seed, and adding the twin does not change them. A fresh retrain was not run for this release, so byte identity of a retrain on another machine or torch build is not verified here.

Provenance

  • Record: shade repository, branch genesis-language, research/cognition/genesis-language/ at 6bf599e25 (the owner's publish decision). Generation I was frozen at b4e92f4ee, with results complete at ca29a1ae0. Synthesis: 2e6d63de7.
  • Study: GL-003c, preregistered at 4d1ea60ee with no amendments; harness at execution commit 9f603776f5186a10438aa2d2a0c87e9fc7e4fa46; decision 54d3d3a87; report 3dddaacd7. The run took 4 h 10 min and 39.0 CPU-hours, and instrument checks I1-I8 were clear.
  • Weights (SHA256SUMS; each file 2,761,077 bytes):
    • seed 29093641: it160 0d8be93ddb60bdd31a2af42a63874d674bf3494a8e3f9d49068d82efe6157d3e, it320 91a76f3a5218eb9fce201486704e87dbbfd560fda0b529ee5fc2e1344410537a, it640 66e10e3f62717cf794efc77c412f74689122e7b6a5d351d8745299f9e4db164e
    • seed 29093642: it160 249ac6d36dfae33ba74c0069ede9be940a252326ec5694261152763eeaae34c4, it320 3e6b73beacc70083a73fef4178370abf07201ca81dbd74c871f3a182267ced4e, it640 cdd05e5df6cdda41952d7ef3e460f502db4e20c03fd73552a41ae76a510ba7bb
    • seed 29093643: it160 223dd9cae3cf126ef0c8d55556cd3a5ecf72fb3adf6b3fbea5d0ccbd26f05d48, it320 ab8d4834ddc98fc8f9c13d68711b59275867b1475214261e41245d1d9e53f978, it640 0f22c4ef38fc242b88dbef3c262cec47d8d9dc1073d7836a4fc0393e48c01779
    • All nine match the record's runs/GL-003c/checkpoint_sha256.txt, checked when this release was staged.
  • Code: code/gl003_dsl.py, code/gl003_learn.py and code/gl003b_learn.py, byte-identical to the record at 6bf599e25. Their SHA-256 are in GENESIS.json.
  • Data: synthetic tasks only, generated by gl003_tasks.py / gl003c_tasks.py from fixed seeds. No external dataset, no pretrained weights and no human text.
  • Host: AMD Ryzen Threadripper 9980X, CPU only, shared with other work under co-tenant load.

Xyntetik-Blad is this house's family name for experimental and edge models. This is a new work trained from scratch, with no parent.