1. Why this blog exists

    The useful part of an experiment usually happens after the screenshot.

  2. I tried to run GLM-5.2 on a 64GB Mac

    Field notes from an experimental ds4 fork, a 244GB GGUF, and the small horror of sparse models that are sparse in compute but not very friendly to filesystems.

  3. Re-quantizing a local model, 14× faster

    Where a 2-bit model spends its bits, and why trying answers used to cost eighty minutes