<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Home on Andrea Borio</title><link>https://andreabor.io/</link><description>Recent content in Home on Andrea Borio</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Fri, 21 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://andreabor.io/index.xml" rel="self" type="application/rss+xml"/><item><title>Now</title><link>https://andreabor.io/now/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0000</pubDate><guid>https://andreabor.io/now/</guid><description>&lt;p&gt;Current work, August 2026:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Qualifying large MoE models in &lt;a href="https://github.com/andreaborio/hebrus"&gt;Hebrus&lt;/a&gt;, with Metal and the SSD treated as one memory system.&lt;/li&gt;&#10;&lt;li&gt;Turning the GLM-5.2 experiment on a 64GB Mac from “does it run?” into the less flattering question: where does the loader waste time?&lt;/li&gt;&#10;&lt;li&gt;Making quantization experiments cheaper with &lt;a href="https://github.com/andreaborio/forgequant"&gt;forgequant&lt;/a&gt;, so a small recipe change does not require rebuilding 1,328 tensors.&lt;/li&gt;&#10;&lt;li&gt;Writing up results only after the second run has had a chance to disagree with the first.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Listening: techno and breakbeat. Recommendations welcome.&lt;/p&gt;</description></item><item><title>Why this blog exists</title><link>https://andreabor.io/posts/why-this-blog-exists/</link><pubDate>Fri, 21 Aug 2026 00:00:00 +0200</pubDate><guid>https://andreabor.io/posts/why-this-blog-exists/</guid><description>&lt;p&gt;The useful part of an experiment usually happens after the screenshot.&lt;/p&gt;&#10;&lt;p&gt;A model runs once. The number looks good. Then the second run changes it, the cache turns out to be warm, or the comparison was not paired. The result becomes less exciting and more useful.&lt;/p&gt;</description></item><item><title>I tried to run GLM-5.2 on a 64GB Mac</title><link>https://andreabor.io/posts/i-tried-to-run-glm-52-on-a-64gb-mac/</link><pubDate>Tue, 23 Jun 2026 14:23:30 +0000</pubDate><guid>https://andreabor.io/posts/i-tried-to-run-glm-52-on-a-64gb-mac/</guid><description>&lt;p&gt;I have a weakness for local LLM experiments that sound slightly unreasonable when said out loud.&lt;/p&gt;&#10;&lt;p&gt;This one was: can GLM-5.2, a very large sparse MoE model, be made to run on a 64GB Apple Silicon Mac without turning the machine into a swap-powered space heater?&lt;/p&gt;</description></item><item><title>Re-quantizing a local model, 14× faster</title><link>https://andreabor.io/posts/re-quantizing-a-local-model-14-faster/</link><pubDate>Wed, 10 Jun 2026 12:49:28 +0000</pubDate><guid>https://andreabor.io/posts/re-quantizing-a-local-model-14-faster/</guid><description>&lt;p&gt;I’m writing this with DeepSeek-V4-Flash running on my Mac, on a coding build I quantized myself. It took about eighty minutes to make. The reason for this post is that I can now rebuild a tweaked version of it in five.&lt;/p&gt;</description></item><item><title>About</title><link>https://andreabor.io/about/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://andreabor.io/about/</guid><description>&lt;p&gt;I&amp;rsquo;m Andrea, a software developer working on local LLM inference.&lt;/p&gt;&#10;&lt;p&gt;I build &lt;a href="https://github.com/andreaborio/hebrus"&gt;Hebrus&lt;/a&gt;, a Metal-first inference engine for Apple Silicon with adaptive SSD streaming. Most of the work starts with a model that does not fit, a machine that should not be enough, and a measurement I do not trust until I can reproduce it.&lt;/p&gt;</description></item></channel></rss>