Skip to content

Latest commit

 

History

742 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ember

rust ci license

Ember runs GGUF on CPU in a way you can inspect and replay. Capture states, patch them, and pack the run into a bundle you can verify later.

Inspectable GGUF forward pass → artifacts with checksums → replayable probes and interventions.

Ember is not a llama.cpp replacement. llama.cpp remains the external performance and correctness reference; Ember focuses on making the states, interventions, and evidence behind a model's output inspectable.

Golden path: Llama-3.2-1B-Instruct Q8_0, with the model and matching tokenizer.json pinned by SHA-256 in the example specs. Other execution paths have different validation coverage.

five-minute workflow

After obtaining the pinned model and tokenizer, put both in the repository root. Ember does not download them automatically. Build once, then capture and verify a run (build and download time are additional):

cargo build --release

target/release/ember experiment validate \
  examples/experiments/morphology-layerwise-capture.toml

target/release/ember experiment run \
  examples/experiments/morphology-layerwise-capture.toml

target/release/ember experiment verify runs/morphology-baseline

The result is runs/morphology-baseline: a verifiable bundle containing provenance and hidden-state captures across all 16 layers for two token selections. A mismatched model or tokenizer fails the pinned identity check.

Patch a state and compare the result:

target/release/ember experiment run \
  examples/experiments/morphology-intervention.toml

target/release/ember experiment compare \
  runs/morphology-baseline runs/morphology-intervention

The full workflow adds restoration and reproduction. On the reference machine, restoration compares bit-exact and baseline reproduction reports exact-semantic. This is a workflow example, not a reproduction of a paper result or a promise of cross-machine bit identity.

why this instrument exists

A probe can read a feature that the model's answer path does not use; more fitting can even make transfer worse. That distinction motivates capture, controlled interventions, and replayable evidence. Read the probe can read it. can the model use it? for the completed eight-ID diagnostic and its limits.

research result: quantization-boundary localization

A deterministic validation wave across Qwen2.5-1.5B and Llama-3.2-1B at Q8, Q6, and Q4 found no evidence of Arabic-selective quantization degradation in the tested matrix.

The surviving result is methodological: Ember can localize rare quantization-boundary failures causally. In validated cases, a single-layer activation patch restored the quantized output, with the causal layer preceding the visible divergence ramp. The observed mechanism was a near-threshold decision flip rather than broad representational collapse.

Model Layers Causal locus
Qwen2.5-1.5B 28 L7
Llama-3.2-1B 16 L1

Validated on the qwen3/llama rows with completed golden checks; see docs/validation.md for the full record.

core documentation

The 1.0 priority is capture, intervention, bundles, and verification on a short, explicitly validated model list. The project remains pre-1.0; an implemented execution path does not imply completed numerical validation.

related tools and research

These are optional extensions and separate research tracks around the CLI instrument:

  • Experiment consoles: native and browser interfaces to the same bundle workflow. The native GUI is experimental and requires Vulkan plus X11/Wayland; it has no software-rendering fallback.
  • Agent runtime: tool protocols and auditable traces, with their own tool-validation and approval boundaries.
  • EmberSEC: hostile-artifact and quantized-fault research. The frozen Phase I evaluation retains its corpus, harnesses, results, and hashes as a separate evidence record; it is not a security certification of the evolving engine. Phase V lab note: Finite Is Not Intact: the Inf result was a fixture; the operational failure is silent finite drift. The measured drift is from synthetic kernels, not a model accuracy drop.
  • Python bindings, maintenance audits, and external benchmarks
  • Sarf Atlas: the separate Arabic morphology workflow package.

citation and license

See CITATION.cff and the MIT license.

About

Inspectable Rust CPU inference for hidden-state capture, causal intervention, and reproducible GGUF experiments.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages