Current work, August 2026:

  • Qualifying large MoE models in Hebrus, with Metal and the SSD treated as one memory system.
  • Turning the GLM-5.2 experiment on a 64GB Mac from “does it run?” into the less flattering question: where does the loader waste time?
  • Making quantization experiments cheaper with forgequant, so a small recipe change does not require rebuilding 1,328 tensors.
  • Writing up results only after the second run has had a chance to disagree with the first.

Listening: techno and breakbeat. Recommendations welcome.