<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Field Notes — AI and engineering</title><description>The main line of Field Notes: Evaluations, task books and pipelines — plus the numbers worth keeping, from model prices to page weight. The other four sections are on the site and in the full feed.</description><link>https://notebookfield.com/</link><language>en</language><item><title>The Chinese open-weights release timeline, dated</title><link>https://notebookfield.com/posts/chinese-open-weights-timeline/</link><guid isPermaLink="true">https://notebookfield.com/posts/chinese-open-weights-timeline/</guid><description>Sixty-seven open-weight releases from Chinese labs, each with a date I could trace to a primary source — plus the eleven candidates I dropped, and why.</description><pubDate>Fri, 09 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;This is a list of dates, not an argument. Every row is a model whose weights were
published for anyone to download, with a date I could trace back to the organization that
published it, a parameter count as stated by that organization, and a link.&lt;/p&gt;
&lt;p&gt;Most timelines of this kind are written from memory or from coverage, which is why they drift.
I built this one from vendor pages, papers and Hugging Face’s API. This revision leaves the table
at 67 rows: seven are new, one of them because an entry I dropped in the first pass turns out to
be open-weights after all. The dropped list at the end has eleven entries, each with a reason.&lt;/p&gt;
&lt;h2 id=&quot;how-to-read-this-timeline&quot;&gt;How to read this timeline&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Open weights only.&lt;/strong&gt; A release is in the table only if the organization published
downloadable weights. API-only and hosted-only models are out, including flagship ones:
Qwen3-Max, the served ERNIE models, Kimi’s hosted variants, DeepSeek’s API-only models.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A date I could trace.&lt;/strong&gt; The Date column has two kinds of value. A plain date is the date on
the organization’s own dated page — a research blog post, a news post, a release page, an
accepted paper. A date marked &lt;code&gt;†&lt;/code&gt; is the Hugging Face repository creation timestamp in UTC,
read from the Hugging Face API on 2026-10-09, used when the vendor had no dated page I could
read. The two can differ by a few days, and that is usually preparation rather than a
different release: DeepSeek-V3’s weights repository was created 2024-12-25 12:52 UTC and the
announcement is dated 2024-12-26.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Parameter counts come from the vendor, or they are blank.&lt;/strong&gt; Total parameters with activated
parameters for mixture-of-experts models, as stated on the announcement, the model card, or
the repository holding the same weights. A family that ships many sizes gets a range. Where
none of the vendor’s pages for that release states a count, the cell is &lt;code&gt;—&lt;/code&gt; rather than a
number carried over from a sibling model, even when the lineage makes the number obvious.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;One link per row, to the vendor or the weight repository.&lt;/strong&gt; No aggregator links, no
news coverage. That makes most rows single-source by construction: each cites the vendor’s own
statement about its own release, and every &lt;code&gt;†&lt;/code&gt; row is single-source by definition.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scope.&lt;/strong&gt; Language models and the flagship multimodal lines that carry the same brand and
weights: Qwen3-VL, Qwen3-Omni, MiniMax-M3. Pure image, video and speech models are out of scope.&lt;/p&gt;
&lt;h2 id=&quot;when-the-vendors-page-is-client-rendered&quot;&gt;When the vendor’s page is client-rendered&lt;/h2&gt;
&lt;p&gt;Two of the labs whose releases I date most often publish announcements as a page that ships no
text to a plain HTTP request. z.ai’s blog returns a 598-byte HTML shell that loads a JavaScript
bundle; qwen.ai’s blog returns a shell that fetches its content after load. Both render fine in
a browser, which is what a reader clicking the link sees.&lt;/p&gt;
&lt;p&gt;For z.ai I did read something: the date string sits inside the page’s own bundle, so the bundle
is where I checked &lt;a href=&quot;https://z.ai/blog/glm-4.5&quot;&gt;GLM-4.5&lt;/a&gt; (2025-07-28),
&lt;a href=&quot;https://z.ai/blog/glm-4.6&quot;&gt;GLM-4.6&lt;/a&gt; (2025-09-30), &lt;a href=&quot;https://z.ai/blog/glm-4.7&quot;&gt;GLM-4.7&lt;/a&gt;
(2025-12-22) and &lt;a href=&quot;https://z.ai/blog/glm-5&quot;&gt;GLM-5&lt;/a&gt; (2026-02-12). That is weaker evidence than a
rendered page, and it is why &lt;a href=&quot;https://z.ai/blog/glm-5.2&quot;&gt;GLM-5.2&lt;/a&gt; stays out: its bundle contains
2026-06-16 and no other date, and a date visible only inside a build artifact is not one I will
print as sourced.&lt;/p&gt;
&lt;p&gt;Qwen’s older blog carries the full text and a date; the newer qwen.ai copies do not. So the
three Qwen3-* rows dated from qwen.ai in the first pass now cite the weight repository and
carry a &lt;code&gt;†&lt;/code&gt;.&lt;/p&gt;
&lt;h2 id=&quot;2023&quot;&gt;2023&lt;/h2&gt;
&lt;h3 id=&quot;q1q2&quot;&gt;Q1–Q2&lt;/h3&gt;
&lt;div class=&quot;table-wrap&quot;&gt;




























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Date&lt;/th&gt;&lt;th&gt;Parameters&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;ChatGLM-6B&lt;/td&gt;&lt;td&gt;2023-03-13 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;6.2B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/THUDM/chatglm-6b&quot;&gt;THUDM/chatglm-6b&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Baichuan-7B&lt;/td&gt;&lt;td&gt;2023-06-13 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;7B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/baichuan-inc/Baichuan-7B&quot;&gt;baichuan-inc/Baichuan-7B&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;ChatGLM2-6B&lt;/td&gt;&lt;td&gt;2023-06-24 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/THUDM/chatglm2-6b&quot;&gt;THUDM/chatglm2-6b&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;h3 id=&quot;q3q4&quot;&gt;Q3–Q4&lt;/h3&gt;
&lt;div class=&quot;table-wrap&quot;&gt;














































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Date&lt;/th&gt;&lt;th&gt;Parameters&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Qwen-7B&lt;/td&gt;&lt;td&gt;2023-08-03 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;7B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/Qwen/Qwen-7B-Chat&quot;&gt;Qwen/Qwen-7B-Chat&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Baichuan 2&lt;/td&gt;&lt;td&gt;2023-09-19&lt;/td&gt;&lt;td&gt;7B, 13B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://arxiv.org/abs/2309.10305&quot;&gt;arXiv:2309.10305&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;ChatGLM3-6B&lt;/td&gt;&lt;td&gt;2023-10-25 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/THUDM/chatglm3-6b&quot;&gt;THUDM/chatglm3-6b&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Yi-34B&lt;/td&gt;&lt;td&gt;2023-11-01 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;34B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/01-ai/Yi-34B&quot;&gt;01-ai/Yi-34B&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek Coder&lt;/td&gt;&lt;td&gt;2023-11-01 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;1.3B–33B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/deepseek-ai/deepseek-coder-33b-instruct&quot;&gt;deepseek-ai/deepseek-coder-33b-instruct&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek LLM&lt;/td&gt;&lt;td&gt;2023-11-29 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;7B, 67B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/deepseek-ai/DeepSeek-LLM-67B-Chat&quot;&gt;deepseek-ai/DeepSeek-LLM-67B-Chat&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;Baichuan 2 is the only 2023 row dated by a paper: &lt;code&gt;arXiv:2309.10305&lt;/code&gt; is dated 2023-09-19,
while its weight repositories appear earlier, on 2023-08-30 &lt;code&gt;†&lt;/code&gt;, so the paper is the single
source for that date. The two ChatGLM cells are blank because their cards do not state a count:
6.2 billion is on the ChatGLM-6B card and not on the cards of its two successors, and copying
it across is the habit that makes timelines of this kind drift.&lt;/p&gt;
&lt;h2 id=&quot;2024&quot;&gt;2024&lt;/h2&gt;
&lt;h3 id=&quot;q1&quot;&gt;Q1&lt;/h3&gt;
&lt;div class=&quot;table-wrap&quot;&gt;
















&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Date&lt;/th&gt;&lt;th&gt;Parameters&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Qwen1.5&lt;/td&gt;&lt;td&gt;2024-02-04&lt;/td&gt;&lt;td&gt;0.5B–110B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://qwenlm.github.io/blog/qwen1.5/&quot;&gt;Qwen blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;Qwen1.5 was in the dropped list in the first pass, because I looked for it on the new blog and
found nothing readable. The old blog still carries the post, dated 2024-02-04, listing six
sizes plus 110B.&lt;/p&gt;
&lt;h3 id=&quot;q2q3&quot;&gt;Q2–Q3&lt;/h3&gt;
&lt;div class=&quot;table-wrap&quot;&gt;














































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Date&lt;/th&gt;&lt;th&gt;Parameters&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-V2&lt;/td&gt;&lt;td&gt;2024-05-07&lt;/td&gt;&lt;td&gt;236B total, 21B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://arxiv.org/abs/2405.04434&quot;&gt;arXiv:2405.04434&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;GLM-4-9B&lt;/td&gt;&lt;td&gt;2024-06-04 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;9B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/zai-org/GLM-4-9B-Chat&quot;&gt;zai-org/GLM-4-9B-Chat&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Qwen2&lt;/td&gt;&lt;td&gt;2024-06-07&lt;/td&gt;&lt;td&gt;0.5B–72B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://qwenlm.github.io/blog/qwen2/&quot;&gt;Qwen blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Qwen2-VL&lt;/td&gt;&lt;td&gt;2024-08-29&lt;/td&gt;&lt;td&gt;2B, 7B, 72B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://qwenlm.github.io/blog/qwen2-vl/&quot;&gt;Qwen blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-V2.5&lt;/td&gt;&lt;td&gt;2024-09-05 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/deepseek-ai/DeepSeek-V2.5&quot;&gt;deepseek-ai/DeepSeek-V2.5&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Qwen2.5&lt;/td&gt;&lt;td&gt;2024-09-19&lt;/td&gt;&lt;td&gt;0.5B–72B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://qwenlm.github.io/blog/qwen2.5/&quot;&gt;Qwen blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;h3 id=&quot;q4&quot;&gt;Q4&lt;/h3&gt;
&lt;div class=&quot;table-wrap&quot;&gt;


































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Date&lt;/th&gt;&lt;th&gt;Parameters&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Hunyuan-Large&lt;/td&gt;&lt;td&gt;2024-10-22 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;389B total, 52B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/tencent/Hunyuan-Large&quot;&gt;tencent/Hunyuan-Large&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Qwen2.5-Coder&lt;/td&gt;&lt;td&gt;2024-11-06 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;0.5B–32B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/Qwen/Qwen2.5-Coder-32B-Instruct&quot;&gt;Qwen/Qwen2.5-Coder-32B-Instruct&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-V2.5-1210&lt;/td&gt;&lt;td&gt;2024-12-10&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://api-docs.deepseek.com/news/news1210&quot;&gt;DeepSeek changelog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-V3&lt;/td&gt;&lt;td&gt;2024-12-26&lt;/td&gt;&lt;td&gt;671B total, 37B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://www.deepseek.com/en/news/deepseek-v3/&quot;&gt;DeepSeek news&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The DeepSeek point releases inside a model line are dated on DeepSeek’s own changelog, one page
per release. Hunyuan-Large’s 389B/52B comes from
&lt;a href=&quot;https://arxiv.org/abs/2411.02265&quot;&gt;arXiv:2411.02265&lt;/a&gt;, submitted 2024-11-04, three weeks after
the weights appeared.&lt;/p&gt;
&lt;h2 id=&quot;2025&quot;&gt;2025&lt;/h2&gt;
&lt;h3 id=&quot;q1-1&quot;&gt;Q1&lt;/h3&gt;
&lt;div class=&quot;table-wrap&quot;&gt;








































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Date&lt;/th&gt;&lt;th&gt;Parameters&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;MiniMax-Text-01&lt;/td&gt;&lt;td&gt;2025-01-14&lt;/td&gt;&lt;td&gt;456B total, 45.9B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://arxiv.org/abs/2501.08313&quot;&gt;arXiv:2501.08313&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-R1&lt;/td&gt;&lt;td&gt;2025-01-20&lt;/td&gt;&lt;td&gt;671B total, 37B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://www.deepseek.com/en/news/deepseek-r1/&quot;&gt;DeepSeek news&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Qwen2.5-VL&lt;/td&gt;&lt;td&gt;2025-01-27 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;3B, 7B, 72B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/Qwen/Qwen2.5-VL-72B-Instruct&quot;&gt;Qwen/Qwen2.5-VL-72B-Instruct&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;QwQ-32B&lt;/td&gt;&lt;td&gt;2025-03-06&lt;/td&gt;&lt;td&gt;32B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://qwenlm.github.io/blog/qwq-32b/&quot;&gt;Qwen blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-V3-0324&lt;/td&gt;&lt;td&gt;2025-03-25&lt;/td&gt;&lt;td&gt;671B total, 37B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://api-docs.deepseek.com/news/news250325&quot;&gt;DeepSeek changelog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The two DeepSeek updates here do not restate a parameter count on their own changelog pages;
the cells carry the count from the repositories that hold those weights, which are the V3 and
R1 repositories listed above.&lt;/p&gt;
&lt;h3 id=&quot;q2&quot;&gt;Q2&lt;/h3&gt;
&lt;div class=&quot;table-wrap&quot;&gt;














































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Date&lt;/th&gt;&lt;th&gt;Parameters&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Qwen3&lt;/td&gt;&lt;td&gt;2025-04-29&lt;/td&gt;&lt;td&gt;0.6B–235B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://qwenlm.github.io/blog/qwen3/&quot;&gt;Qwen blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-R1-0528&lt;/td&gt;&lt;td&gt;2025-05-28&lt;/td&gt;&lt;td&gt;671B total, 37B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://api-docs.deepseek.com/news/news250528&quot;&gt;DeepSeek changelog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MiniCPM4-8B&lt;/td&gt;&lt;td&gt;2025-06-05 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;8B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/openbmb/MiniCPM4-8B&quot;&gt;openbmb/MiniCPM4-8B&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MiniMax-M1&lt;/td&gt;&lt;td&gt;2025-06-16&lt;/td&gt;&lt;td&gt;456B total, 45.9B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://arxiv.org/abs/2506.13585&quot;&gt;arXiv:2506.13585&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Hunyuan-A13B&lt;/td&gt;&lt;td&gt;2025-06-27&lt;/td&gt;&lt;td&gt;80B total, 13B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://github.com/Tencent-Hunyuan/Hunyuan-A13B&quot;&gt;Tencent-Hunyuan/Hunyuan-A13B&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;ERNIE 4.5&lt;/td&gt;&lt;td&gt;2025-06-30&lt;/td&gt;&lt;td&gt;0.3B–300B; flagship 300B/47B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://ernie.baidu.com/blog/posts/ernie4.5/&quot;&gt;ERNIE blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;h3 id=&quot;q3&quot;&gt;Q3&lt;/h3&gt;
&lt;div class=&quot;table-wrap&quot;&gt;
























































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Date&lt;/th&gt;&lt;th&gt;Parameters&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Kimi K2&lt;/td&gt;&lt;td&gt;2025-07-11 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;1T total, 32B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/moonshotai/Kimi-K2-Instruct&quot;&gt;moonshotai/Kimi-K2-Instruct&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Qwen3-Coder&lt;/td&gt;&lt;td&gt;2025-07-22&lt;/td&gt;&lt;td&gt;480B total, 35B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://qwenlm.github.io/blog/qwen3-coder/&quot;&gt;Qwen blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;GLM-4.5&lt;/td&gt;&lt;td&gt;2025-07-28&lt;/td&gt;&lt;td&gt;355B total, 32B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://z.ai/blog/glm-4.5&quot;&gt;z.ai blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Step-3&lt;/td&gt;&lt;td&gt;2025-07-31&lt;/td&gt;&lt;td&gt;321B total, 38B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://chat.stepfun.com/research/zh/step3&quot;&gt;StepFun research&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Seed-OSS-36B&lt;/td&gt;&lt;td&gt;2025-08-21&lt;/td&gt;&lt;td&gt;36B&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://seed.bytedance.com/en/blog/seed-oss-open-source-models-release&quot;&gt;ByteDance Seed blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-V3.1&lt;/td&gt;&lt;td&gt;2025-08-21 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/deepseek-ai/DeepSeek-V3.1&quot;&gt;deepseek-ai/DeepSeek-V3.1&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;LongCat-Flash-Chat&lt;/td&gt;&lt;td&gt;2025-08-29 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;560B total, 18.6B–31.3B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/meituan-longcat/LongCat-Flash-Chat&quot;&gt;meituan-longcat/LongCat-Flash-Chat&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Qwen3-Next-80B-A3B&lt;/td&gt;&lt;td&gt;2025-09-09 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;80B total, 3B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct&quot;&gt;Qwen/Qwen3-Next-80B-A3B-Instruct&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Ling-flash-2.0&lt;/td&gt;&lt;td&gt;2025-09-17 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;100B total, 6.1B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/inclusionAI/Ling-flash-2.0&quot;&gt;inclusionAI/Ling-flash-2.0&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Qwen3-Omni-30B-A3B&lt;/td&gt;&lt;td&gt;2025-09-20 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;30B total, 3B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct&quot;&gt;Qwen/Qwen3-Omni-30B-A3B-Instruct&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Qwen3-VL&lt;/td&gt;&lt;td&gt;2025-09-22 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;235B total, 22B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/Qwen/Qwen3-VL-235B-A22B-Instruct&quot;&gt;Qwen/Qwen3-VL-235B-A22B-Instruct&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-V3.2-Exp&lt;/td&gt;&lt;td&gt;2025-09-29&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://www.deepseek.com/en/news/v3-2-exp/&quot;&gt;DeepSeek news&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;GLM-4.6&lt;/td&gt;&lt;td&gt;2025-09-30&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://z.ai/blog/glm-4.6&quot;&gt;z.ai blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;h3 id=&quot;q4-1&quot;&gt;Q4&lt;/h3&gt;
&lt;div class=&quot;table-wrap&quot;&gt;




















































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Date&lt;/th&gt;&lt;th&gt;Parameters&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Ling-1T&lt;/td&gt;&lt;td&gt;2025-10-02 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;1T total, 50B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/inclusionAI/Ling-1T&quot;&gt;inclusionAI/Ling-1T&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-OCR&lt;/td&gt;&lt;td&gt;2025-10-17 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/deepseek-ai/DeepSeek-OCR&quot;&gt;deepseek-ai/DeepSeek-OCR&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;LongCat-Flash-Omni&lt;/td&gt;&lt;td&gt;2025-10-23 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;560B total, 27B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/meituan-longcat/LongCat-Flash-Omni&quot;&gt;meituan-longcat/LongCat-Flash-Omni&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MiniMax-M2&lt;/td&gt;&lt;td&gt;2025-10-27&lt;/td&gt;&lt;td&gt;230B total, 10B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://www.minimax.io/blog/minimax-m2-en-1748600000&quot;&gt;MiniMax blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Kimi K2 Thinking&lt;/td&gt;&lt;td&gt;2025-11-04 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/moonshotai/Kimi-K2-Thinking&quot;&gt;moonshotai/Kimi-K2-Thinking&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-V3.2&lt;/td&gt;&lt;td&gt;2025-12-01&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://www.deepseek.com/en/news/deepseek-v3-2/&quot;&gt;DeepSeek news&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;GLM-4.7&lt;/td&gt;&lt;td&gt;2025-12-22&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://z.ai/blog/glm-4.7&quot;&gt;z.ai blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;h2 id=&quot;2026&quot;&gt;2026&lt;/h2&gt;
&lt;h3 id=&quot;q1-2&quot;&gt;Q1&lt;/h3&gt;
&lt;div class=&quot;table-wrap&quot;&gt;








































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Date&lt;/th&gt;&lt;th&gt;Parameters&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;GLM-4.7-Flash&lt;/td&gt;&lt;td&gt;2026-01-19 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;30B total, 3B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/zai-org/GLM-4.7-Flash&quot;&gt;zai-org/GLM-4.7-Flash&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Step-3.5-Flash&lt;/td&gt;&lt;td&gt;2026-02-01 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;196B total, 11B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/stepfun-ai/Step-3.5-Flash&quot;&gt;stepfun-ai/Step-3.5-Flash&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MiniMax-M2.5&lt;/td&gt;&lt;td&gt;2026-02-12&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://www.minimax.io/news/minimax-m25&quot;&gt;MiniMax news&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;GLM-5&lt;/td&gt;&lt;td&gt;2026-02-12&lt;/td&gt;&lt;td&gt;744B total, 40B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://z.ai/blog/glm-5&quot;&gt;z.ai blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Qwen3.5-397B-A17B&lt;/td&gt;&lt;td&gt;2026-02-16&lt;/td&gt;&lt;td&gt;397B total, 17B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://home.alibabagroup.com/en-US/document-1960233590314762240&quot;&gt;Alibaba press release&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;h3 id=&quot;q2-1&quot;&gt;Q2&lt;/h3&gt;
&lt;div class=&quot;table-wrap&quot;&gt;








































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Date&lt;/th&gt;&lt;th&gt;Parameters&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;MiniMax-M2.7&lt;/td&gt;&lt;td&gt;2026-04-09 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/MiniMaxAI/MiniMax-M2.7&quot;&gt;MiniMaxAI/MiniMax-M2.7&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Kimi K2.6&lt;/td&gt;&lt;td&gt;2026-04-14 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;1T total, 32B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/moonshotai/Kimi-K2.6&quot;&gt;moonshotai/Kimi-K2.6&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-V4 (Pro and Flash)&lt;/td&gt;&lt;td&gt;2026-04-24&lt;/td&gt;&lt;td&gt;Pro 1.6T total / 49B active; Flash 284B total / 13B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://www.deepseek.com/en/news/v4-preview&quot;&gt;DeepSeek news&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Step 3.7 Flash&lt;/td&gt;&lt;td&gt;2026-05-23 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;198B total, 11B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/stepfun-ai/Step-3.7-Flash&quot;&gt;stepfun-ai/Step-3.7-Flash&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MiniMax-M3&lt;/td&gt;&lt;td&gt;2026-06-01&lt;/td&gt;&lt;td&gt;~428B total, ~23B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://www.minimax.io/blog/minimax-m3&quot;&gt;MiniMax blog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The DeepSeek-V4 row now carries both counts; the Flash model’s 284B total is stated on the same
announcement page as the Pro figures. Step 3.7 Flash moved to a &lt;code&gt;†&lt;/code&gt; date: its GitHub repository
has no releases, and the page I first cited carries a creation timestamp rather than a release
date.&lt;/p&gt;
&lt;h3 id=&quot;q3-1&quot;&gt;Q3&lt;/h3&gt;
&lt;div class=&quot;table-wrap&quot;&gt;














































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Date&lt;/th&gt;&lt;th&gt;Parameters&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;LongCat-2.0&lt;/td&gt;&lt;td&gt;2026-07-05 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;1.6T total, ~48B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/meituan-longcat/LongCat-2.0&quot;&gt;meituan-longcat/LongCat-2.0&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Hunyuan Hy3&lt;/td&gt;&lt;td&gt;2026-07-06&lt;/td&gt;&lt;td&gt;295B total, 21B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://www.tencent.com/tencent-hunyuan-officially-releases-hy3-advancing-agent-capabilities-and-deeper-product-integration/&quot;&gt;Tencent newsroom&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Qwen3.8&lt;/td&gt;&lt;td&gt;2026-08-05 &lt;code&gt;†&lt;/code&gt;&lt;/td&gt;&lt;td&gt;27B; 2.4T total / 95B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B&quot;&gt;Qwen/Qwen3.8-2.4T-A95B&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-V4-Pro GA&lt;/td&gt;&lt;td&gt;2026-08-13&lt;/td&gt;&lt;td&gt;1.6T total, 49B active&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://api-docs.deepseek.com/news/news260813&quot;&gt;DeepSeek changelog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-V4-Flash-Vision-Exp&lt;/td&gt;&lt;td&gt;2026-08-21&lt;/td&gt;&lt;td&gt;&lt;code&gt;—&lt;/code&gt;&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://api-docs.deepseek.com/news/news260821&quot;&gt;DeepSeek changelog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek-V4.1-Flash&lt;/td&gt;&lt;td&gt;2026-09-10&lt;/td&gt;&lt;td&gt;552B total; 8B active for input, 16B for output&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://api-docs.deepseek.com/news/news260910&quot;&gt;DeepSeek changelog&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;Hy3 is the entry whose status changed in this revision. In the first pass I dropped it because
the page I found did not say whether weights were published. It does: Hy3 is open-sourced under
Apache 2.0 with weights on Hugging Face, and Tencent’s newsroom names the parameter counts. The
V4-Pro GA row repeats the count from the April preview announcement, because the GA page does
not restate it. The Qwen3.8 row is dated by the repository of the 27B model, the first of the
line to appear; the flagship 2.4T-A95B repository was created three days later, and Qwen’s own
announcement page is one of the client-rendered ones described above.&lt;/p&gt;
&lt;p&gt;Two rows elsewhere disagree with third-party timelines for the same reason: LongCat-2.0’s
repository was created 2026-07-05 while secondary coverage puts its launch at the end of June,
and DeepSeek-V4-Flash’s repository was created 2026-04-22, two days before the announcement
page.&lt;/p&gt;
&lt;h2 id=&quot;what-changed-in-this-revision&quot;&gt;What changed in this revision&lt;/h2&gt;
&lt;p&gt;Seven rows were added and three cells changed, listed because the edit history is the point of
a table like this:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Qwen1.5 promoted into 2024 Q1&lt;/strong&gt; (2024-02-04), after finding the dated post on the old
Qwen blog.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Hunyuan Hy3 promoted into 2026 Q3&lt;/strong&gt; (2026-07-06). It was dropped in error.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Four DeepSeek point releases added&lt;/strong&gt;: V2.5-1210 (2024-12-10), V3-0324 (2025-03-25),
R1-0528 (2025-05-28) and V4-Pro GA (2026-08-13). All four are dated on DeepSeek’s own
changelog, the same source as V3.2 and V3.2-Exp.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Qwen3.8 added&lt;/strong&gt; (2026-08-05 &lt;code&gt;†&lt;/code&gt;), which the first pass omitted entirely.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Two ChatGLM parameter cells blanked.&lt;/strong&gt; The ChatGLM2-6B and ChatGLM3-6B cards do not state
a count, so the cells are now &lt;code&gt;—&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;DeepSeek-V4-Flash gained its total.&lt;/strong&gt; The row said “Flash 13B active” and dropped the
284B total, which the same page states.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Three rows changed source or date type for the same structural reason: the Qwen3-Next,
Qwen3-Omni and Step 3.7 Flash rows now cite a weight repository with a &lt;code&gt;†&lt;/code&gt;, because the pages
cited in the first pass were either unreadable to a plain request or carried no release date.&lt;/p&gt;
&lt;h2 id=&quot;how-this-table-gets-updated&quot;&gt;How this table gets updated&lt;/h2&gt;
&lt;p&gt;The table is edited in place. Every quarter I re-read the vendor pages for the labs already
listed, then check Hugging Face for new repositories under those organizations. A release
enters the table when it has a date and a link; until then it sits in the dropped list, with
the reason.&lt;/p&gt;
&lt;p&gt;Two rules keep it stable: I do not replace a &lt;code&gt;†&lt;/code&gt; date with a later one unless the vendor
publishes its own dated page, and I do not remove a superseded row.&lt;/p&gt;
&lt;p&gt;The current table has 67 entries. Twenty-three of them fall in the twelve months from October
2025 to October 2026. That is a count of this table and nothing more; it is not a claim about
the field, because the table only holds releases I could date.&lt;/p&gt;
&lt;h2 id=&quot;what-i-could-not-check&quot;&gt;What I could not check&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Eleven candidates were dropped&lt;/strong&gt;, for these reasons:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Yi-1.5&lt;/strong&gt; (May 2024) — the weight repository
(&lt;a href=&quot;https://huggingface.co/01-ai/Yi-1.5-34B&quot;&gt;01-ai/Yi-1.5-34B&lt;/a&gt;) was created 2024-05-11 &lt;code&gt;†&lt;/code&gt;, and
I could not date the announcement itself.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;InternLM2 and InternLM2.5&lt;/strong&gt; — the weight repositories
(&lt;a href=&quot;https://huggingface.co/internlm/internlm2-7b&quot;&gt;internlm/internlm2-7b&lt;/a&gt;,
&lt;a href=&quot;https://huggingface.co/internlm/internlm2_5-7b-chat&quot;&gt;internlm/internlm2_5-7b-chat&lt;/a&gt;) look
like renamed predecessors, so the creation timestamp does not mark the release.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;MiniCPM 1 and 2&lt;/strong&gt; — same problem; the repository
(&lt;a href=&quot;https://huggingface.co/openbmb/MiniCPM-2B-sft-bf16&quot;&gt;openbmb/MiniCPM-2B-sft-bf16&lt;/a&gt;) is now a
monorepo for a whole series.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Baichuan 3 and 4&lt;/strong&gt; — I could not find weight artifacts at all.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ERNIE 5.0&lt;/strong&gt; and &lt;strong&gt;ERNIE 5.1&lt;/strong&gt; — Baidu’s blog dates them 2026-02-06 and 2026-05-09
(&lt;a href=&quot;https://ernie.baidu.com/blog/posts/ernie5.0/&quot;&gt;ERNIE 5.0&lt;/a&gt;,
&lt;a href=&quot;https://ernie.baidu.com/blog/posts/ernie-5.1-0508-release/&quot;&gt;ERNIE 5.1&lt;/a&gt;) and describes
ERNIE 5.0 as a 2.4-trillion-parameter model, but I found no open-weight repository for
either.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Kimi K3&lt;/strong&gt; — the model card (&lt;a href=&quot;https://huggingface.co/moonshotai/Kimi-K3&quot;&gt;moonshotai/Kimi-K3&lt;/a&gt;)
carries the full architecture table (2.8T total, 104B activated) and a license, but no
release date. The repository timestamp is 2026-06-13 and secondary coverage says mid-July. I
could not reconcile that.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Kimi K2.5&lt;/strong&gt; — same shape of problem: the weight repository timestamp
(&lt;a href=&quot;https://huggingface.co/moonshotai/Kimi-K2.5&quot;&gt;moonshotai/Kimi-K2.5&lt;/a&gt;) is 2026-01-01 while the
launch is reported at the end of January.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GLM-5.2&lt;/strong&gt; — the post is client-rendered and the only date string in its bundle is
2026-06-16. I am not willing to print that as the release date.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;DeepSeek-V3.1-Terminus&lt;/strong&gt; — a mid-cycle rename with a 2025-09-22 &lt;code&gt;†&lt;/code&gt; repository
(&lt;a href=&quot;https://huggingface.co/deepseek-ai/DeepSeek-V3.1-Terminus&quot;&gt;deepseek-ai/DeepSeek-V3.1-Terminus&lt;/a&gt;);
including it would have double-counted the V3.1 line.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Qwen-Image-2.1&lt;/strong&gt; — a dated weight repository
(&lt;a href=&quot;https://huggingface.co/Qwen/Qwen-Image-2.1&quot;&gt;Qwen/Qwen-Image-2.1&lt;/a&gt;, 2026-09-14 &lt;code&gt;†&lt;/code&gt;) but an
image model, outside this table’s scope by the rule stated at the top.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The Step-Audio line&lt;/strong&gt; — the same, for speech. The repositories exist
(&lt;a href=&quot;https://huggingface.co/stepfun-ai/Step-Audio-R1.1&quot;&gt;stepfun-ai/Step-Audio-R1.1&lt;/a&gt;) and I did
not date them, because they would need their own table with its own scope.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Three other limits on this table:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The &lt;code&gt;†&lt;/code&gt; timestamps are preparation times, not launch times.&lt;/strong&gt; A repository can be created
days before the announcement, and from outside I cannot see whether weights were uploaded
at once or in pieces. Where the gap is large I noted it above the table.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The dates are the vendor’s own.&lt;/strong&gt; No release date here has been confirmed by a second
source, and one link per release is deliberate: a vendor correcting itself is the signal
worth keeping, and coverage repeating the vendor is not a second source.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;License differences are not reflected.&lt;/strong&gt; Apache-2.0, MIT and the vendor-specific
licenses Kimi, MiniMax and LongCat ship under are not the same permission. Each row links to
the page that states its license, and for Hy3 that page states Apache 2.0.&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Measuring error in support AI, and what accuracy hides</title><link>https://notebookfield.com/posts/hallucination-rates-in-support-ai/</link><guid isPermaLink="true">https://notebookfield.com/posts/hallucination-rates-in-support-ai/</guid><description>Five kinds of error, three ways one accuracy figure misleads, and what a routing test on 3,894 public complaints actually returned.</description><pubDate>Fri, 09 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The way I test a support or ticketing system is against a task book, not a demo. The number people
ask for afterwards is accuracy, and I have stopped handing that over on its own: on every realistic
test set I have built, the average was the least informative thing in the results file.&lt;/p&gt;
&lt;p&gt;What follows is the method I use now, plus a run on public data so the numbers can be checked.&lt;/p&gt;
&lt;h2 id=&quot;five-failure-modes-five-different-bills&quot;&gt;Five failure modes, five different bills&lt;/h2&gt;
&lt;p&gt;A support model fails in at least five ways, and they do not cost the same.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fabrication.&lt;/strong&gt; The reply asserts something absent from the policy or the record: a fee that does
not exist, a wrong refund window, an invented delivery date. This is what people call
hallucination. The taxonomy I lean on, OlaBench, splits it into four sub-types — factual
hallucination, misuse of retrieved results, relevance hallucination, and logical inconsistency
(&lt;a href=&quot;https://arxiv.org/abs/2510.22143v3&quot;&gt;OlaBench, arXiv:2510.22143v3&lt;/a&gt;, §3.1; v1 25 Oct 2025, v3 25
May 2026, retrieved 9 Oct 2026). A sentence that reads well and is false costs money, and the cost
usually arrives months later.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Missed detection.&lt;/strong&gt; The ticket contains a legal threat, a regulator’s name, a safety issue or a
fraud signal, and the system files it as routine. Nothing false was said. The failure is that no
human ever reads it. On a scoring sheet this often counts as a correct classification, because the
model did pick a queue that exists.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Misrouting.&lt;/strong&gt; The reply is reasonable and the queue is wrong, so a differently-skilled team
applies its own policy. The measurable cost is time to resolution and repeat contact.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Out-of-scope answers.&lt;/strong&gt; The system answers a question it was never built for — tax, legal
advice, another provider’s product. The content may be accurate. The exposure is regulatory.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Refusal.&lt;/strong&gt; The system declines, loops, or hands off. There is no answer, so there is no error to
score. This one gets tracked as a transfer or containment rate and drops out of the error count
entirely.&lt;/p&gt;
&lt;div class=&quot;table-wrap&quot;&gt;








































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Failure&lt;/th&gt;&lt;th&gt;What is wrong&lt;/th&gt;&lt;th&gt;Where the cost lands&lt;/th&gt;&lt;th&gt;Commonly reported as&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Fabrication&lt;/td&gt;&lt;td&gt;the content&lt;/td&gt;&lt;td&gt;refunds, goodwill, disputes&lt;/td&gt;&lt;td&gt;hallucination rate&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Missed detection&lt;/td&gt;&lt;td&gt;nothing&lt;/td&gt;&lt;td&gt;escalation, legal, safety&lt;/td&gt;&lt;td&gt;almost nothing&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Misrouting&lt;/td&gt;&lt;td&gt;the destination&lt;/td&gt;&lt;td&gt;time to resolution, repeats&lt;/td&gt;&lt;td&gt;accuracy&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Out of scope&lt;/td&gt;&lt;td&gt;the remit&lt;/td&gt;&lt;td&gt;compliance&lt;/td&gt;&lt;td&gt;rarely measured&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Refusal&lt;/td&gt;&lt;td&gt;nothing&lt;/td&gt;&lt;td&gt;churn, containment&lt;/td&gt;&lt;td&gt;transfer rate&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;One accuracy figure blends these five at whatever ratio the sample happens to contain, and that
ratio is set by whoever picked the sample.&lt;/p&gt;
&lt;h2 id=&quot;what-i-ran&quot;&gt;What I ran&lt;/h2&gt;
&lt;p&gt;I needed a routing task with a ground-truth label and text a model would really see. Public
complaint data has both.&lt;/p&gt;
&lt;p&gt;The US CFPB publishes its Consumer Complaint Database under CC0 (&lt;a href=&quot;https://www.consumerfinance.gov/data-research/consumer-complaints/&quot;&gt;CFPB Consumer Complaint
Database&lt;/a&gt;, retrieved 9 Oct
2026); the search API’s own metadata reports 18,274,023 records, license CC0, last updated
2026-10-08 (&lt;a href=&quot;https://www.consumerfinance.gov/data-research/consumer-complaints/search/api/v1/?size=0&amp;#x26;no_aggs=true&quot;&gt;API metadata&lt;/a&gt;,
retrieved 9 Oct 2026).&lt;/p&gt;
&lt;p&gt;The JSON search endpoint returns 15 fields and no narrative (&lt;a href=&quot;https://cfpb.github.io/api/ccdb/fields.html&quot;&gt;field
reference&lt;/a&gt;, retrieved 9 Oct 2026). That is new: the
September 2026 release removed complaint narratives from the database (&lt;a href=&quot;https://cfpb.github.io/api/ccdb/release-notes.html&quot;&gt;CFPB release
notes&lt;/a&gt;, Release 24, September 2026). So I took
the text from a Hugging Face mirror instead, &lt;code&gt;BEE-spoke-data/consumer-finance-complaints&lt;/code&gt; (&lt;a href=&quot;https://huggingface.co/datasets/BEE-spoke-data/consumer-finance-complaints&quot;&gt;dataset
card&lt;/a&gt;, CC0, mirror last
modified 29 Dec 2025); its &lt;code&gt;has-text&lt;/code&gt; config holds 1,689,573 rows (&lt;a href=&quot;https://datasets-server.huggingface.co/size?dataset=BEE-spoke-data%2Fconsumer-finance-complaints&quot;&gt;split
sizes&lt;/a&gt;,
retrieved 9 Oct 2026).&lt;/p&gt;
&lt;p&gt;Task: read the narrative only, predict the product queue. Ground truth: the product label the CFPB
publishes. I drew 3,894 complaints from it. The oldest is dated 2015-03-20, the newest
2024-01-04.&lt;/p&gt;
&lt;h2 id=&quot;the-label-set-changed-under-the-data&quot;&gt;The label set changed under the data&lt;/h2&gt;
&lt;p&gt;The raw product field holds 17 distinct strings in this sample, and they are not 17 products. The
date ranges are what give it away.&lt;/p&gt;
&lt;div class=&quot;table-wrap&quot;&gt;






























































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Product string as published&lt;/th&gt;&lt;th&gt;Rows&lt;/th&gt;&lt;th&gt;Received, first → last in this sample&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Credit reporting, credit repair services, or other personal consumer reports&lt;/td&gt;&lt;td&gt;1,892&lt;/td&gt;&lt;td&gt;2017-04-24 → 2023-08-24&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Debt collection&lt;/td&gt;&lt;td&gt;547&lt;/td&gt;&lt;td&gt;2015-03-20 → 2023-12-18&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Credit card or prepaid card&lt;/td&gt;&lt;td&gt;267&lt;/td&gt;&lt;td&gt;2017-04-24 → 2023-08-21&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mortgage&lt;/td&gt;&lt;td&gt;248&lt;/td&gt;&lt;td&gt;2015-04-21 → 2023-11-02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Credit reporting or other personal consumer reports&lt;/td&gt;&lt;td&gt;245&lt;/td&gt;&lt;td&gt;2023-08-26 → 2024-01-04&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Checking or savings account&lt;/td&gt;&lt;td&gt;195&lt;/td&gt;&lt;td&gt;2017-04-24 → 2023-11-21&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Credit reporting&lt;/td&gt;&lt;td&gt;92&lt;/td&gt;&lt;td&gt;2015-04-17 → 2017-04-09&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Student loan&lt;/td&gt;&lt;td&gt;85&lt;/td&gt;&lt;td&gt;2015-10-06 → 2023-11-13&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Money transfer, virtual currency, or money service&lt;/td&gt;&lt;td&gt;79&lt;/td&gt;&lt;td&gt;2017-08-08 → 2023-11-13&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Vehicle loan or lease&lt;/td&gt;&lt;td&gt;75&lt;/td&gt;&lt;td&gt;2017-04-24 → 2023-11-01&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Credit card&lt;/td&gt;&lt;td&gt;68&lt;/td&gt;&lt;td&gt;2015-04-06 → 2023-12-22&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Payday loan, title loan, or personal loan&lt;/td&gt;&lt;td&gt;41&lt;/td&gt;&lt;td&gt;2017-04-24 → 2023-06-20&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Bank account or service&lt;/td&gt;&lt;td&gt;34&lt;/td&gt;&lt;td&gt;2015-04-04 → 2017-01-31&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Consumer Loan&lt;/td&gt;&lt;td&gt;19&lt;/td&gt;&lt;td&gt;2015-03-25 → 2017-02-03&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Payday loan&lt;/td&gt;&lt;td&gt;3&lt;/td&gt;&lt;td&gt;2015-10-08 → 2017-02-24&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Money transfers&lt;/td&gt;&lt;td&gt;2&lt;/td&gt;&lt;td&gt;2015-07-17 → 2015-12-05&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Payday loan, title loan, personal loan, or advance loan&lt;/td&gt;&lt;td&gt;2&lt;/td&gt;&lt;td&gt;2023-09-20 → 2023-10-06&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;Read the dates down the table. &lt;code&gt;Bank account or service&lt;/code&gt; last appears here on 2017-01-31, and
&lt;code&gt;Checking or savings account&lt;/code&gt; first appears on 2017-04-24. Three more strings end in 2017 and their
successors start the same year. The three credit-reporting strings sit end to end: 2015 to 2017,
2017 to 2023, 2023 into 2024. Same category, renamed.&lt;/p&gt;
&lt;p&gt;I collapsed the 17 strings into 9 categories before scoring anything. The mapping is in the script,
applied in one place rather than inferred per row. Without that step an accuracy compares a 2016
row and a 2023 row whose label sets do not line up, and the model gets credit or blame for the
taxonomy.&lt;/p&gt;
&lt;p&gt;The live database disagrees with the mirror about which names are current. Aggregated over
complaints received in 2026, the CFPB API returns 11 product strings, with &lt;code&gt;Credit card&lt;/code&gt; and
&lt;code&gt;Prepaid card&lt;/code&gt; listed separately and no &lt;code&gt;Credit card or prepaid card&lt;/code&gt; at all (&lt;a href=&quot;https://www.consumerfinance.gov/data-research/consumer-complaints/search/api/v1/?size=0&amp;#x26;aggs=product&amp;#x26;date_received_min=2026-01-01&quot;&gt;product aggregation,
complaints received 2026-01-01
onward&lt;/a&gt;,
retrieved 9 Oct 2026). I did not reconcile the two taxonomies.&lt;/p&gt;
&lt;h2 id=&quot;three-ways-one-accuracy-figure-misleads&quot;&gt;Three ways one accuracy figure misleads&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Imbalance.&lt;/strong&gt; After collapsing, the sample looks like this:&lt;/p&gt;
&lt;div class=&quot;table-wrap&quot;&gt;






















































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Collapsed category&lt;/th&gt;&lt;th&gt;Rows&lt;/th&gt;&lt;th&gt;Share&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Credit reporting&lt;/td&gt;&lt;td&gt;2,229&lt;/td&gt;&lt;td&gt;57.2%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Debt collection&lt;/td&gt;&lt;td&gt;547&lt;/td&gt;&lt;td&gt;14.0%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Credit card or prepaid card&lt;/td&gt;&lt;td&gt;335&lt;/td&gt;&lt;td&gt;8.6%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mortgage&lt;/td&gt;&lt;td&gt;248&lt;/td&gt;&lt;td&gt;6.4%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Checking or savings account&lt;/td&gt;&lt;td&gt;229&lt;/td&gt;&lt;td&gt;5.9%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Student loan&lt;/td&gt;&lt;td&gt;85&lt;/td&gt;&lt;td&gt;2.2%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Money transfer or virtual currency&lt;/td&gt;&lt;td&gt;81&lt;/td&gt;&lt;td&gt;2.1%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Vehicle loan or lease&lt;/td&gt;&lt;td&gt;75&lt;/td&gt;&lt;td&gt;1.9%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Personal / payday / title loan&lt;/td&gt;&lt;td&gt;65&lt;/td&gt;&lt;td&gt;1.7%&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;A predictor that answers “credit reporting” for every ticket scores 0.572 here. That is a
constant, and it sits close enough to a real model’s score to be mistaken for one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sample selection.&lt;/strong&gt; Same data, same model, two splits:&lt;/p&gt;
&lt;div class=&quot;table-wrap&quot;&gt;

























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Split&lt;/th&gt;&lt;th&gt;Test n&lt;/th&gt;&lt;th&gt;Accuracy&lt;/th&gt;&lt;th&gt;Macro F1&lt;/th&gt;&lt;th&gt;Credit-reporting share of test set&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Random 60/40&lt;/td&gt;&lt;td&gt;1,558&lt;/td&gt;&lt;td&gt;0.756&lt;/td&gt;&lt;td&gt;0.428&lt;/td&gt;&lt;td&gt;57.1%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Chronological, split at 2022-04-21&lt;/td&gt;&lt;td&gt;1,560&lt;/td&gt;&lt;td&gt;0.832&lt;/td&gt;&lt;td&gt;0.377&lt;/td&gt;&lt;td&gt;74.0%&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The chronological run scores 7.6 points higher while getting worse on most categories. The test
set became more concentrated in the biggest class, and the dominant-class predictor rides that
concentration up. Anyone reporting the 0.832 as an improvement after retraining would be reading
their own sampling as progress.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Coverage.&lt;/strong&gt; Sort the test set by the model’s confidence and score only the top slice:&lt;/p&gt;
&lt;div class=&quot;table-wrap&quot;&gt;


































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Coverage&lt;/th&gt;&lt;th&gt;Tickets scored&lt;/th&gt;&lt;th&gt;Accuracy on that slice&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;100%&lt;/td&gt;&lt;td&gt;1,558&lt;/td&gt;&lt;td&gt;0.756&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;75%&lt;/td&gt;&lt;td&gt;1,168&lt;/td&gt;&lt;td&gt;0.854&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;50%&lt;/td&gt;&lt;td&gt;779&lt;/td&gt;&lt;td&gt;0.904&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;25%&lt;/td&gt;&lt;td&gt;389&lt;/td&gt;&lt;td&gt;0.900&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;10%&lt;/td&gt;&lt;td&gt;155&lt;/td&gt;&lt;td&gt;0.897&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The 0.904 at half coverage is a real number computed on held-out data, and published alone it
misdescribes the system by 15 points. Note where it stops moving: from 50% coverage down to 10%,
accuracy drifts from 0.904 to 0.897, so the model’s confidence and its correctness come apart on
the tail. That tail is where the misroutes are.&lt;/p&gt;
&lt;p&gt;Underneath the average, the per-category picture on the random split:&lt;/p&gt;
&lt;div class=&quot;table-wrap&quot;&gt;
































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Category&lt;/th&gt;&lt;th&gt;Test n&lt;/th&gt;&lt;th&gt;Recall&lt;/th&gt;&lt;th&gt;Precision&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Credit reporting&lt;/td&gt;&lt;td&gt;889&lt;/td&gt;&lt;td&gt;0.879&lt;/td&gt;&lt;td&gt;0.886&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Debt collection&lt;/td&gt;&lt;td&gt;225&lt;/td&gt;&lt;td&gt;0.631&lt;/td&gt;&lt;td&gt;0.637&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Credit card or prepaid card&lt;/td&gt;&lt;td&gt;129&lt;/td&gt;&lt;td&gt;0.760&lt;/td&gt;&lt;td&gt;0.430&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mortgage&lt;/td&gt;&lt;td&gt;114&lt;/td&gt;&lt;td&gt;0.860&lt;/td&gt;&lt;td&gt;0.737&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Checking or savings account&lt;/td&gt;&lt;td&gt;93&lt;/td&gt;&lt;td&gt;0.548&lt;/td&gt;&lt;td&gt;0.607&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Money transfer or virtual currency&lt;/td&gt;&lt;td&gt;32&lt;/td&gt;&lt;td&gt;0.000&lt;/td&gt;&lt;td&gt;0.000&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Student loan&lt;/td&gt;&lt;td&gt;29&lt;/td&gt;&lt;td&gt;0.276&lt;/td&gt;&lt;td&gt;0.889&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Vehicle loan or lease&lt;/td&gt;&lt;td&gt;29&lt;/td&gt;&lt;td&gt;0.000&lt;/td&gt;&lt;td&gt;0.000&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Personal / payday / title loan&lt;/td&gt;&lt;td&gt;18&lt;/td&gt;&lt;td&gt;0.000&lt;/td&gt;&lt;td&gt;0.000&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;Three categories are never once predicted correctly. Accuracy is 0.756. Macro F1 is 0.428. The
distance between those two is the honest headline for this model.&lt;/p&gt;
&lt;h2 id=&quot;pulling-the-high-risk-errors-out-of-the-average&quot;&gt;Pulling the high-risk errors out of the average&lt;/h2&gt;
&lt;p&gt;The useful published idea here is OlaBench’s risk axis, Critical Business Risk Rate. It counts only
responses that inappropriately assert one of four things, each a high-stakes failure that may
trigger compliance exposure, user disputes or reputational damage: admitting platform liability
(618 test cases), misidentifying the ICS role (228), overcommitting (138), and disparaging
individuals or merchants (16) (&lt;a href=&quot;https://arxiv.org/abs/2510.22143v3&quot;&gt;OlaBench,
arXiv:2510.22143v3&lt;/a&gt;, §3.1 and Table 1; 25 May 2026, retrieved 9
Oct 2026). It is reported as its own number, beside the hallucination rate.&lt;/p&gt;
&lt;p&gt;The shape of that is what to copy: choose failures whose consequences differ in kind, count them
separately, publish both, and never average them.&lt;/p&gt;
&lt;p&gt;For the routing test I defined three tiers, and I chose them:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;critical&lt;/strong&gt; — a complaint about funds or account access (checking or savings, money transfer)
lands in the reporting queue&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;material&lt;/strong&gt; — a credit or loan dispute lands in a different loan or reporting family&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;routine&lt;/strong&gt; — a misroute inside the same family&lt;/li&gt;
&lt;/ul&gt;
&lt;div class=&quot;table-wrap&quot;&gt;
























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Tier&lt;/th&gt;&lt;th&gt;Random split&lt;/th&gt;&lt;th&gt;Chronological split&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Critical&lt;/td&gt;&lt;td&gt;7 / 1,558 = 0.4%&lt;/td&gt;&lt;td&gt;11 / 1,560 = 0.7%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Material&lt;/td&gt;&lt;td&gt;43 / 1,558 = 2.8%&lt;/td&gt;&lt;td&gt;28 / 1,560 = 1.8%&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Routine&lt;/td&gt;&lt;td&gt;330 / 1,558 = 21.2%&lt;/td&gt;&lt;td&gt;223 / 1,560 = 14.3%&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The two splits swap places depending on which tier you read: the one with the prettier accuracy
hides a critical error rate almost twice as high.&lt;/p&gt;
&lt;p&gt;Published work also shows how far the same named metric can move when the population changes. The
OlaBench paper reports a critical business risk rate of 8.7% for its OlaMind-Stage-2 model on the
benchmark, and a critical business risk rate below 0.05% in daily manual annotation of live
dialogues after launch (&lt;a href=&quot;https://arxiv.org/abs/2510.22143v3&quot;&gt;OlaBench, arXiv:2510.22143v3&lt;/a&gt;, §5.2 and
Table 3; 25 May 2026, retrieved 9 Oct 2026). Both are the same metric, measured on different
populations with different annotators. Quoting one without the other is a choice.&lt;/p&gt;
&lt;h2 id=&quot;how-these-numbers-were-produced&quot;&gt;How these numbers were produced&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Data.&lt;/strong&gt; 3,894 complaints from a CC0 mirror of the CFPB Consumer Complaint Database (config
&lt;code&gt;has-text&lt;/code&gt;, 1,689,573 rows). The draw is 40 pages of 100 consecutive rows at random offsets in
the split, deduplicated by complaint ID, seed 20261009. Rows with a narrative under 80 characters
were dropped. It is a clustered draw, not an independent one.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Label hygiene.&lt;/strong&gt; 17 published product strings collapsed to 9 categories using the date ranges
in the table above. The mapping is written out in the script, not inferred per row.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Model.&lt;/strong&gt; Multinomial naive Bayes over word unigrams and bigrams, Laplace smoothing α = 1,
features seen at least twice in training. Python standard library only.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Splits.&lt;/strong&gt; Random 60/40 with seed 7; chronological split at the 60th percentile of the
received-date order, which lands on 2022-04-21.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Selective accuracy.&lt;/strong&gt; Test rows sorted by the model’s maximum posterior probability, top k
scored.&lt;/li&gt;
&lt;li&gt;Every number in the tables above is printed by that one script over that sample.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deliberate choice.&lt;/strong&gt; No language model appears in this pipeline, including the scoring. A
measurement instrument that is itself a model carries its own error, and OlaBench publishes the
size of its own: its judge agrees with human annotators at 91.7% on risk identification and 82.6%
on hallucination detection, over 5,000 human-annotated instances (&lt;a href=&quot;https://arxiv.org/abs/2510.22143v3&quot;&gt;OlaBench,
arXiv:2510.22143v3&lt;/a&gt;, §3.3 and Table 2; 25 May 2026, retrieved
9 Oct 2026). A judge that disagrees with people 8% of the time on the risk axis puts a floor under
everything else you can measure with it.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;what-i-could-not-check&quot;&gt;What I could not check&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Whether the mirror’s product strings match the live database’s current taxonomy. The mirror’s
rows use names the live API no longer returns for complaints received in 2026. I collapsed by date
range and did not reconcile the two.&lt;/li&gt;
&lt;li&gt;Whether the mirror is complete. Its files are timestamped December 2025 and the newest complaint
in my 3,894-row sample is dated 2024-01-04. I did not verify whether later records exist in it.&lt;/li&gt;
&lt;li&gt;Whether complaints published with a narrative are representative of all complaints. Narratives
appeared only with consent, so every number above sits on that consenting subset.&lt;/li&gt;
&lt;li&gt;Whether this sample’s class mix reflects the population. Credit reporting is 57.2% of the sample
and about 82% of the live database: 14,978,722 of 18,274,023 records carry one of the three
credit-reporting strings (&lt;a href=&quot;https://www.consumerfinance.gov/data-research/consumer-complaints/search/api/v1/?size=0&amp;#x26;aggs=product&quot;&gt;product
aggregation&lt;/a&gt;,
retrieved 9 Oct 2026). Most of the gap is the mirror’s date cap.&lt;/li&gt;
&lt;li&gt;How noisy the product label is. It is chosen by the consumer at submission, not assigned by an
expert, and I treated it as ground truth because companies route on it. If it is noisy, my error
rates are a floor, not a ceiling.&lt;/li&gt;
&lt;li&gt;Whether a published tiering exists for complaint-routing consequences. I could not find one, so
the three tiers are mine. The OlaBench categories are a different kind of thing: they score what
a reply asserts about liability and commitment, and they do not classify misroutes.&lt;/li&gt;
&lt;li&gt;Whether error rates vary with company size, state, or complaint age. A stratified rate is the
next thing to compute.&lt;/li&gt;
&lt;li&gt;The naive Bayes model is a floor. A trained transformer would score higher on the same task. The
run was to see the shape of the error distribution, and a better model changes the level without
changing the shape.&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>How an eval set is built, and why it rots</title><link>https://notebookfield.com/posts/how-an-eval-set-is-built/</link><guid isPermaLink="true">https://notebookfield.com/posts/how-an-eval-set-is-built/</guid><description>An eval set has a purpose, a scoring rule, a calibration record and an expiry date. Where samples come from, what counts as correct, what retires it.</description><pubDate>Fri, 09 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;An eval set is an instrument. It has a purpose, a scoring rule, a calibration record and an expiry date. Treating it as a file you download is where the trouble starts: the file already made those decisions without writing them down.&lt;/p&gt;
&lt;p&gt;What follows is the build document I would want for a set I have to defend six months from now. Each rule is traceable to a public source or marked as my practice.&lt;/p&gt;
&lt;h2 id=&quot;where-the-samples-come-from&quot;&gt;Where the samples come from&lt;/h2&gt;
&lt;p&gt;Four sources, each importing a different bias.&lt;/p&gt;
&lt;div class=&quot;table-wrap&quot;&gt;


































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Source&lt;/th&gt;&lt;th&gt;What you get&lt;/th&gt;&lt;th&gt;The bias it imports&lt;/th&gt;&lt;th&gt;Documented where&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;A public benchmark, reused&lt;/td&gt;&lt;td&gt;Comparability with published numbers&lt;/td&gt;&lt;td&gt;Contamination; a task choice someone else made&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://arxiv.org/abs/2310.18018&quot;&gt;Sainz et al., 2023-10-27&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Public administrative data&lt;/td&gt;&lt;td&gt;A real distribution, a public licence, dated updates&lt;/td&gt;&lt;td&gt;A sample of who reports, not of who is affected&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://cfpb.github.io/api/ccdb/index.html&quot;&gt;CFPB, read 2026-10-09&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Synthetic generation&lt;/td&gt;&lt;td&gt;Volume you can afford, labels you control&lt;/td&gt;&lt;td&gt;Generator artifacts in the surface form&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://aclanthology.org/N18-2017.pdf&quot;&gt;Gururangan et al., ACL 2018&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Your own production traffic&lt;/td&gt;&lt;td&gt;Items that match what you ship&lt;/td&gt;&lt;td&gt;A convenience sample that drifts with the product&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://arxiv.org/abs/1802.03916&quot;&gt;Lipton et al., ICML 2018&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Reuse.&lt;/strong&gt; LiveBench’s authors state the problem in one sentence: test set contamination “can quickly render benchmarks obsolete” (&lt;a href=&quot;https://arxiv.org/abs/2406.19314&quot;&gt;LiveBench, v2 2025-04-18&lt;/a&gt;). Their answer was questions from recent competitions, arXiv papers and news, added monthly. LiveCodeBench did the same for code: 400 problems published between May 2023 and February 2024 (&lt;a href=&quot;https://arxiv.org/abs/2403.07974&quot;&gt;LiveCodeBench, arXiv v1 2024-03-12&lt;/a&gt;). A static benchmark decays without anyone editing it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Administrative data.&lt;/strong&gt; The CFPB complaint database is the best example I know. Its search API reported 18,274,023 records under a CC0 licence on 2026-10-09, with a &lt;code&gt;last_updated&lt;/code&gt; of 2026-10-08 (&lt;a href=&quot;https://www.consumerfinance.gov/data-research/consumer-complaints/search/api/v1/?size=0&amp;#x26;no_aggs=true&quot;&gt;CFPB complaint search API, read 2026-10-09&lt;/a&gt;). Complaints are published only after the company responds or after 15 days, whichever comes first, and complaints referred to other regulators, including depository institutions with less than $10 billion in assets, are not published (&lt;a href=&quot;https://cfpb.github.io/api/ccdb/index.html&quot;&gt;CFPB API docs, read 2026-10-09&lt;/a&gt;). The agency states that the database “is not a statistical sample of consumers’ experiences in the marketplace” (&lt;a href=&quot;https://www.consumerfinance.gov/data-research/consumer-complaints/&quot;&gt;CFPB database page, read 2026-10-09&lt;/a&gt;). A random draw samples the people who complained and whose complaints survived publication, not the people who were harmed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Synthetic.&lt;/strong&gt; Gururangan et al. showed that a text classifier reading only the hypothesis, with the premise deleted, reached about 67% on SNLI and 53% on MultiNLI (&lt;a href=&quot;https://aclanthology.org/N18-2017.pdf&quot;&gt;Annotation Artifacts in NLI Data, ACL 2018&lt;/a&gt;). The label was recoverable from phrasing because the crowd workers followed a template. Every generator has one. My test: delete the part of the input meant to carry the answer, score again, and compare with the majority-class rate. If accuracy stays well above it, the set is partly measuring its own generator. &lt;em&gt;(my practice)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Production sampling.&lt;/strong&gt; This is the source I trust least and use most. It produces items that match what the system sees, and its composition changes when the product does. Label shift is the named version of that failure: the marginal &lt;code&gt;p(y)&lt;/code&gt; moves while &lt;code&gt;p(x|y)&lt;/code&gt; does not, measurable without test labels using Black Box Shift Estimation (&lt;a href=&quot;https://arxiv.org/abs/1802.03916&quot;&gt;Lipton et al., ICML 2018&lt;/a&gt;).&lt;/p&gt;
&lt;h2 id=&quot;what-counts-as-correct&quot;&gt;What counts as correct&lt;/h2&gt;
&lt;p&gt;The match rule is a parameter, and the same predictions score differently under each choice. Four conventions:&lt;/p&gt;
&lt;div class=&quot;table-wrap&quot;&gt;





























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Convention&lt;/th&gt;&lt;th&gt;The rule, as the source writes it&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Exact match against any reference&lt;/td&gt;&lt;td&gt;Counts predictions that match any one ground truth answer exactly; punctuation and articles ignored&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://arxiv.org/abs/1606.05250&quot;&gt;SQuAD, arXiv v3 2016-10-11&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Partial credit against the best reference&lt;/td&gt;&lt;td&gt;F1 over bags of tokens, “the maximum F1 over all of the ground truth answers”, then averaged&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://arxiv.org/abs/1606.05250&quot;&gt;SQuAD, arXiv v3 2016-10-11&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;A threshold on localisation&lt;/td&gt;&lt;td&gt;IoU thresholds &lt;code&gt;[.5:.05:.95]&lt;/code&gt;, 10 of them; 101 recall thresholds; &lt;code&gt;maxDets&lt;/code&gt; of 1, 10 and 100&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://github.com/cocodataset/cocoapi/blob/master/PythonAPI/pycocotools/cocoeval.py&quot;&gt;cocoapi &lt;code&gt;cocoeval.py&lt;/code&gt;, read 2026-10-09&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;A graded rubric with an exclusion flag&lt;/td&gt;&lt;td&gt;0–3 for how well specified the issue is; 0–3 for whether the tests are scoped to it; a separate 0/1 question for any other reason the sample should not be used; a 1–5 confidence rating&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://cdn.openai.com/introducing-swe-bench-verified/swe-b-annotation-instructions.pdf&quot;&gt;SWE-bench annotation instructions, undated&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;Three things this table taught me.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Say who the ceiling belongs to.&lt;/strong&gt; In SQuAD the human number comes from the data: the second annotator’s answer is treated as the prediction, the rest as ground truth, and humans score 77.0 exact match and 86.8 F1 on the test set (&lt;a href=&quot;https://arxiv.org/abs/1606.05250&quot;&gt;SQuAD, arXiv v3 2016-10-11&lt;/a&gt;). A ceiling quoted without its rule says as much about the rule as about the task.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;One published number can hide ten thresholds.&lt;/strong&gt; COCO’s headline detection score averages over ten IoU thresholds, three detection budgets and a recall grid. A single figure without a named threshold is a claim nobody can re-check. Pick the threshold, publish it, keep it fixed across versions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The exclusion question is the part people skip.&lt;/strong&gt; SWE-bench Verified is 500 samples “verified to be non-problematic by our human annotators”, drawn from 1,699 randomly sampled items that 93 Python developers annotated, with the full annotation set and the rubric released alongside (&lt;a href=&quot;https://openai.com/index/introducing-swe-bench-verified/&quot;&gt;OpenAI, updated 2025-02-24&lt;/a&gt;). The rubric’s last section asks, 0/1, whether there is any other reason not to use the sample. My handling of it, &lt;em&gt;(my practice)&lt;/em&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A closed list of exclusion reason codes, decided before annotation starts.&lt;/li&gt;
&lt;li&gt;Two raters must agree before an item is dropped; one rater’s “unclear” is not an exemption.&lt;/li&gt;
&lt;li&gt;The dropped item stays in the file with its reason code, so the exemption stays auditable.&lt;/li&gt;
&lt;li&gt;The exemption rate is reported next to every result; if it moves between releases, the set changed even when the questions did not.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Ground truth built by pooling is biased by construction.&lt;/strong&gt; Test collections are built by judging only a sample of documents. When the pool is small relative to the collection, the judgment set “can be biased in that they favor relevant documents that contain topic title words”, a bias that depends on collection size and not on the number of relevant documents (&lt;a href=&quot;https://www.nist.gov/publications/bias-and-limits-pooling-large-collections&quot;&gt;Buckley et al., NIST, 2007-07-17&lt;/a&gt;). In a modern eval, if your labels come from checking only what your current systems returned, the label set favours items phrased the way those systems phrase things. My rule: draw a random slice from outside the candidate pool. &lt;em&gt;(my practice)&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;whether-the-raters-agree&quot;&gt;Whether the raters agree&lt;/h2&gt;
&lt;p&gt;Agreement is a property of the pair and the setup.&lt;/p&gt;
&lt;p&gt;Krippendorff’s alpha is &lt;code&gt;1 − Do/De&lt;/code&gt;, and it accepts any number of coders and missing entries (&lt;a href=&quot;https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf&quot;&gt;Krippendorff, 2011-01-25&lt;/a&gt;). For two raters, Cohen’s kappa, with Landis and Koch’s 1977 bands (&lt;a href=&quot;https://www.ncbi.nlm.nih.gov/books/NBK52665/table/ch3.t5&quot;&gt;reproduced in an AHRQ report, 2010&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;Here is why the setup has to travel with the number. First-turn results on MT-bench, in the paper’s two setups: S1 counts ties and inconsistent votes, S2 excludes ties (&lt;a href=&quot;https://arxiv.org/abs/2306.05685&quot;&gt;MT-Bench, v4 2023-12-24&lt;/a&gt;):&lt;/p&gt;
&lt;div class=&quot;table-wrap&quot;&gt;




























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Pair (first turn)&lt;/th&gt;&lt;th&gt;S1, R = 33%&lt;/th&gt;&lt;th&gt;S2, R = 50%&lt;/th&gt;&lt;th&gt;Source&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;GPT-4 pairwise vs human expert&lt;/td&gt;&lt;td&gt;66% (n = 1,343)&lt;/td&gt;&lt;td&gt;85% (n = 859)&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://arxiv.org/abs/2306.05685&quot;&gt;MT-Bench, v4 2023-12-24&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;GPT-4 single-answer vs human expert&lt;/td&gt;&lt;td&gt;60% (n = 1,280)&lt;/td&gt;&lt;td&gt;85% (n = 739)&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://arxiv.org/abs/2306.05685&quot;&gt;MT-Bench, v4 2023-12-24&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Human vs human&lt;/td&gt;&lt;td&gt;63% (n = 721)&lt;/td&gt;&lt;td&gt;81% (n = 479)&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://arxiv.org/abs/2306.05685&quot;&gt;MT-Bench, v4 2023-12-24&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The headline everyone quotes is 85% against 81% for human agreement (&lt;a href=&quot;https://arxiv.org/abs/2306.05685&quot;&gt;MT-Bench, v4 2023-12-24&lt;/a&gt;). The same table holds 66% for the same judge under the other tie rule. An agreement figure quoted without its setup is not comparable to anything.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Measure agreement per stratum, not pooled.&lt;/strong&gt; ChaosNLI collected 464,500 annotations, 100 per example, over 3,113 SNLI and MNLI examples and 1,532 ANLI examples (&lt;a href=&quot;https://arxiv.org/abs/2010.03532&quot;&gt;Nie et al., 2020-10-08&lt;/a&gt;). Models are near-perfect where humans agree and “can barely beat a random guess” where they do not, and those low-agreement items are most of the errors. A pooled score mixes two measurement regimes whose mixing ratio changes as models improve.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sizing.&lt;/strong&gt; For an agreement study, start from the kappa minimum-sample-size tables (&lt;a href=&quot;https://riviste.unimi.it/index.php/ebph/article/view/17614&quot;&gt;Bujang and Baharum, EBPH 14(2), 2017&lt;/a&gt;). For comparing two systems, use a power analysis: Miller gives the MDE formula and shows that raising answers per question from 1 to 10 moved the MDE from 13.2% to 7.5%; in his illustrative table, HumanEval’s 164 questions at a fictional 83.6% carry a standard error of 3.2% (&lt;a href=&quot;https://arxiv.org/abs/2411.00640&quot;&gt;Miller, 2024-11-01&lt;/a&gt;). If items arrive in groups, such as the same document or template, compute clustered standard errors (&lt;a href=&quot;https://arxiv.org/abs/2411.00640&quot;&gt;Miller, 2024-11-01&lt;/a&gt;, Table 4); on DROP they were 1.34 against 0.44 naive.&lt;/p&gt;
&lt;p&gt;My rule: size the calibration slice for the agreement interval I need, the scoring slice for the effect I want to detect, and refuse to report a gap below the MDE. &lt;em&gt;(my practice)&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&quot;why-the-set-rots&quot;&gt;Why the set rots&lt;/h2&gt;
&lt;div class=&quot;table-wrap&quot;&gt;





























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Decay&lt;/th&gt;&lt;th&gt;Symptom&lt;/th&gt;&lt;th&gt;What I do&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Contamination&lt;/td&gt;&lt;td&gt;One set improves while unrelated sets do not&lt;/td&gt;&lt;td&gt;Keep item dates and prompt hashes; treat a stale set as historical&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Label shift&lt;/td&gt;&lt;td&gt;Scores move while the system did not&lt;/td&gt;&lt;td&gt;Estimate the label mix and re-weight or re-cut the slice&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Taxonomy and rubric drift&lt;/td&gt;&lt;td&gt;The same question means something different&lt;/td&gt;&lt;td&gt;Version the labels; freeze the rubric text per set version&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Saturation&lt;/td&gt;&lt;td&gt;Every system sits near the ceiling&lt;/td&gt;&lt;td&gt;Retire the set and say when&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;Contamination’s extent “is unknown, as it is not straightforward to measure”, and where it exists it overestimates performance (&lt;a href=&quot;https://arxiv.org/abs/2310.18018&quot;&gt;Sainz et al., 2023-10-27&lt;/a&gt;). I cannot audit training corpora, so I keep the dates and admit the exposure window.&lt;/p&gt;
&lt;p&gt;Saturation is not a failure of the set: GLUE’s performance passed the level of non-expert humans, which is what prompted SuperGLUE (&lt;a href=&quot;https://arxiv.org/abs/1905.00537&quot;&gt;Wang et al., 2019-05-02&lt;/a&gt;). Benchmark choice is fragile on its own: swapping tasks can reorder methods even when the methods do not change (&lt;a href=&quot;https://arxiv.org/abs/2107.07002&quot;&gt;Dehghani et al., 2021-07-14&lt;/a&gt;). Both are reasons to keep more than one set.&lt;/p&gt;
&lt;p&gt;Taxonomy drift has a public paper trail. The CFPB’s September 2026 release notes record that “consumers’ complaint narratives … have been removed from the database”; the June 2026 release removed “Consumer disputed” and “Consumer consent provided” from exports, long after those filters went away; the May 2023 release re-based the ZIP code field on 2019 census estimates (&lt;a href=&quot;https://cfpb.github.io/api/ccdb/release-notes.html&quot;&gt;CFPB release notes, read 2026-10-09&lt;/a&gt;). A set built on those narratives became unmaintainable one month ago. The same page explains the general rule: the database shows “the consumer’s original products, sub-products, issues, and sub-issues selections consistent with the options available on the form at the time the consumer submitted the complaint” (&lt;a href=&quot;https://www.consumerfinance.gov/data-research/consumer-complaints/&quot;&gt;CFPB database page, read 2026-10-09&lt;/a&gt;). Labels frozen at submission, against a taxonomy that keeps moving.&lt;/p&gt;
&lt;h2 id=&quot;keeping-it-maintainable&quot;&gt;Keeping it maintainable&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;(my practice, with borrowed pieces named)&lt;/em&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Version the set.&lt;/strong&gt; Additions and exclusions get a version bump and a changelog line; never edit an item silently.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Two slices.&lt;/strong&gt; A frozen slice for comparability, a rolling window for freshness. LiveCodeBench and LiveBench are public versions of the same idea.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Write the metadata down.&lt;/strong&gt; A datasheet (&lt;a href=&quot;https://arxiv.org/abs/1803.09010&quot;&gt;Gebru et al., 2018, CACM 2021&lt;/a&gt;) or a Croissant record (&lt;a href=&quot;https://arxiv.org/abs/2403.19546&quot;&gt;2024-03-28&lt;/a&gt;), so the next person need not infer the collection process.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A desensitisation rule that is checkable.&lt;/strong&gt; The CFPB’s ZIP rule is a good template: publish the 5-digit ZIP unless the census area has under 20,000 people, then 3 digits only if the wider area has more than 20,000, otherwise nothing (&lt;a href=&quot;https://cfpb.github.io/api/ccdb/fields.html&quot;&gt;CFPB field reference, read 2026-10-09&lt;/a&gt;). A rule stated as a threshold can be re-run; “we removed personal information” cannot.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Watch the exemption rate.&lt;/strong&gt; It is the leading indicator that the rubric stopped fitting the data.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;State the blind spot in one line.&lt;/strong&gt; SQuAD’s metrics ignore punctuation and articles; COCO’s headline averages ten thresholds. Both are fine as long as the limitation is on the label.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Plan the retirement.&lt;/strong&gt; When a set saturates or its taxonomy moves, freeze it, date it, and build the next one from the same seed items.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;what-i-could-not-check&quot;&gt;What I could not check&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;How many SWE-bench samples were dropped to leave 500.&lt;/strong&gt; The announcement gives the annotated total (1,699) and the rubric, not the drop count (&lt;a href=&quot;https://openai.com/index/introducing-swe-bench-verified/&quot;&gt;OpenAI, updated 2025-02-24&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Whether any specific model saw any specific benchmark item.&lt;/strong&gt; Sainz et al. state the extent is unknown; I have no method that resolves this, so I do not claim cleanliness.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The CFPB’s narrative scrubbing standard.&lt;/strong&gt; The field is gone and the current field reference no longer documents it; I could not reach a dated scrubbing specification.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The CFPB’s historical &lt;code&gt;consumer_consent_provided&lt;/code&gt; semantics.&lt;/strong&gt; The export field and its filter are gone; I found the removal dates, not the values.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Krippendorff’s own minimum sample size for alpha.&lt;/strong&gt; I have the kappa tables from Bujang and Baharum; I found no equivalent table for alpha.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The Landis and Koch bands at first hand.&lt;/strong&gt; I read a reproduction of the 1977 table in a 2010 AHRQ report, not the original.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LiveCodeBench’s current item count.&lt;/strong&gt; The 400 figure comes from the March 2024 paper; the platform’s live count may differ.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The OpenAI SWE-bench Verified page at first hand.&lt;/strong&gt; The live page served an automated-client challenge when I read it; the same version is readable in an &lt;a href=&quot;https://web.archive.org/web/20260101233457/https://openai.com/index/introducing-swe-bench-verified/&quot;&gt;Internet Archive copy&lt;/a&gt;, and the annotation instruction PDF is undated.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The CFPB record count above is a live API value&lt;/strong&gt; and changes daily. It was 18,274,023 when I read it on 2026-10-09.&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>From token prices to task prices: three jobs, seven APIs</title><link>https://notebookfield.com/posts/llm-api-pricing-per-task/</link><guid isPermaLink="true">https://notebookfield.com/posts/llm-api-pricing-per-task/</guid><description>I priced three concrete jobs on seven model APIs using each vendor&apos;s own pricing page, and wrote down every assumption the conversion rests on.</description><pubDate>Fri, 09 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Every vendor publishes a price per million tokens. Almost none publishes what a job costs; the number that decides whether a feature ships is left as arithmetic homework.&lt;/p&gt;
&lt;p&gt;I did that arithmetic for three jobs, from first-party price lists read on 2026-10-09. The token counts are assumptions, not measurements, and they are written out below. The money is the vendor’s standard-tier list price, in USD.&lt;/p&gt;
&lt;h2 id=&quot;the-three-jobs&quot;&gt;The three jobs&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Job 1, label 1,000 support messages.&lt;/strong&gt; Each call sends a 400-token instruction block and a 120-token message, and asks for a 15-token answer. Billed per run: 520,000 input tokens, of which 400,000 are the instruction block sent 1,000 times, plus 15,000 output tokens.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Job 2, turn a 100,000-word draft into a structured summary.&lt;/strong&gt; Google’s token guide says a token is about 4 characters and 100 tokens about 60-80 English words, putting 100,000 words at 125,000-167,000 tokens (&lt;a href=&quot;https://ai.google.dev/gemini-api/docs/tokens&quot;&gt;Google token counting guide&lt;/a&gt;, read 2026-10-09). I used 140,000:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;One shot:&lt;/strong&gt; 140,000 tokens in, a 3,000-token structured summary out.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Map-reduce:&lt;/strong&gt; 14 chunk calls of about 10,000 tokens each, 700 tokens of notes per chunk, then a 9,800-token reduce call. Totals: 149,800 in, 12,800 out.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Job 3, one round of tool-assisted bug fixing.&lt;/strong&gt; Eight turns against a 20,000-token repository: system prompt, tool schemas, three files. The prompt grows about 2,500 tokens per turn as tool results come back, from 20,000 to 37,500. Cumulative billed input is 230,000 tokens, of which 192,500 are a re-sent prefix a cache can serve; output is 6,000 tokens including reasoning.&lt;/p&gt;
&lt;h2 id=&quot;how-i-read-the-price-lists&quot;&gt;How I read the price lists&lt;/h2&gt;
&lt;p&gt;Five first-party pages, all read on 2026-10-09; the Gemini API page builds its tables client-side, so I read it in a browser.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI API pricing&lt;/a&gt; — standard tier&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic pricing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;Gemini API pricing&lt;/a&gt; — paid tier, introductory rates through 2026-12-31&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://api-docs.deepseek.com/quick_start/pricing&quot;&gt;DeepSeek Models &amp;#x26; Pricing&lt;/a&gt; — off-peak column&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.x.ai/developers/pricing&quot;&gt;xAI pricing&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The rules I applied:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Standard tier only&lt;/strong&gt;, unless a row says batch: the real-time tier.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cache reads are billed at the published cached-input rate. Cache writes are billed where a vendor publishes a write price.&lt;/strong&gt; OpenAI and Anthropic do: $2.50 per million tokens for the two frontier models here, $0.125 for the two small ones, on Anthropic’s 5-minute cache. Google, DeepSeek and xAI publish no write price, so I charged zero, which flatters those three in every cached row.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;In the cached rows I charge a write for every input token that is not a cache read:&lt;/strong&gt; 400 tokens in Job 1, 37,500 in Job 3.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reasoning tokens are output tokens.&lt;/strong&gt; OpenAI states they are billed as output tokens (&lt;a href=&quot;https://developers.openai.com/api/docs/guides/reasoning&quot;&gt;OpenAI reasoning guide&lt;/a&gt;, read 2026-10-09). Gemini’s output price is quoted “including thinking tokens” and DeepSeek’s V4.1-Flash runs in thinking mode by default.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool and search fees sit outside the main table.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;No prompt hits a long-context threshold&lt;/strong&gt;; the prompt-length step that does apply, Anthropic’s 100,000-token rule, is below.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;No aggregator site was used&lt;/strong&gt; for any price.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;what-each-job-costs&quot;&gt;What each job costs&lt;/h2&gt;
&lt;div class=&quot;table-wrap&quot;&gt;
































































































































































































































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Task&lt;/th&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Estimated cost&lt;/th&gt;&lt;th&gt;Assumptions&lt;/th&gt;&lt;th&gt;Price source (read 2026-10-09)&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Job 1, no cache&lt;/td&gt;&lt;td&gt;gpt-6.1-sol&lt;/td&gt;&lt;td&gt;$1.19&lt;/td&gt;&lt;td&gt;standard tier; 520k in / 15k out&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 1, no cache&lt;/td&gt;&lt;td&gt;gpt-6-luna&lt;/td&gt;&lt;td&gt;$0.060&lt;/td&gt;&lt;td&gt;standard tier; 520k in / 15k out&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 1, no cache&lt;/td&gt;&lt;td&gt;Claude Sonnet 5.5&lt;/td&gt;&lt;td&gt;$1.19&lt;/td&gt;&lt;td&gt;standard tier; 520k in / 15k out&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 1, no cache&lt;/td&gt;&lt;td&gt;Claude Haiku 5.5&lt;/td&gt;&lt;td&gt;$0.060&lt;/td&gt;&lt;td&gt;standard tier; 520k in / 15k out&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 1, no cache&lt;/td&gt;&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;&lt;td&gt;$0.45&lt;/td&gt;&lt;td&gt;standard tier, introductory rate; 520k in / 15k out&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;Google&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 1, no cache&lt;/td&gt;&lt;td&gt;DeepSeek V4.1-Flash&lt;/td&gt;&lt;td&gt;$0.087&lt;/td&gt;&lt;td&gt;off-peak window; 520k in / 15k out&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://api-docs.deepseek.com/quick_start/pricing&quot;&gt;DeepSeek&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 1, no cache&lt;/td&gt;&lt;td&gt;Grok 4.7&lt;/td&gt;&lt;td&gt;$1.13&lt;/td&gt;&lt;td&gt;standard tier; 520k in / 15k out&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://docs.x.ai/developers/pricing&quot;&gt;xAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 1, instructions cached&lt;/td&gt;&lt;td&gt;gpt-6.1-sol&lt;/td&gt;&lt;td&gt;$0.43&lt;/td&gt;&lt;td&gt;400k of input served from cache; one 400-token write&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 1, instructions cached&lt;/td&gt;&lt;td&gt;gpt-6-luna&lt;/td&gt;&lt;td&gt;$0.024&lt;/td&gt;&lt;td&gt;400k of input served from cache; one 400-token write&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 1, instructions cached&lt;/td&gt;&lt;td&gt;Claude Sonnet 5.5&lt;/td&gt;&lt;td&gt;$0.43&lt;/td&gt;&lt;td&gt;400k cached, 5-minute TTL and write&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 1, instructions cached&lt;/td&gt;&lt;td&gt;Claude Haiku 5.5&lt;/td&gt;&lt;td&gt;$0.024&lt;/td&gt;&lt;td&gt;400k cached, prompts under 100k, 5-minute TTL&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 1, instructions cached&lt;/td&gt;&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;&lt;td&gt;$0.18&lt;/td&gt;&lt;td&gt;400k of input served from cache, storage excluded&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;Google&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 1, instructions cached&lt;/td&gt;&lt;td&gt;DeepSeek V4.1-Flash&lt;/td&gt;&lt;td&gt;$0.028&lt;/td&gt;&lt;td&gt;off-peak; cache hits free of a write fee&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://api-docs.deepseek.com/quick_start/pricing&quot;&gt;DeepSeek&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 1, instructions cached&lt;/td&gt;&lt;td&gt;Grok 4.7&lt;/td&gt;&lt;td&gt;$0.53&lt;/td&gt;&lt;td&gt;400k of input served from cache&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://docs.x.ai/developers/pricing&quot;&gt;xAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, one shot&lt;/td&gt;&lt;td&gt;gpt-6.1-sol&lt;/td&gt;&lt;td&gt;$0.31&lt;/td&gt;&lt;td&gt;140k in / 3k out, one call&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, one shot&lt;/td&gt;&lt;td&gt;gpt-6-luna&lt;/td&gt;&lt;td&gt;$0.016&lt;/td&gt;&lt;td&gt;140k in / 3k out, one call&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, one shot&lt;/td&gt;&lt;td&gt;Claude Sonnet 5.5&lt;/td&gt;&lt;td&gt;$0.31&lt;/td&gt;&lt;td&gt;140k in / 3k out, one call&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, one shot&lt;/td&gt;&lt;td&gt;Claude Haiku 5.5&lt;/td&gt;&lt;td&gt;$0.078&lt;/td&gt;&lt;td&gt;140k in / 3k out, above the 100k step&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, one shot&lt;/td&gt;&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;&lt;td&gt;$0.12&lt;/td&gt;&lt;td&gt;140k in / 3k out, one call&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;Google&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, one shot&lt;/td&gt;&lt;td&gt;DeepSeek V4.1-Flash&lt;/td&gt;&lt;td&gt;$0.023&lt;/td&gt;&lt;td&gt;off-peak; 140k in / 3k out&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://api-docs.deepseek.com/quick_start/pricing&quot;&gt;DeepSeek&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, one shot&lt;/td&gt;&lt;td&gt;Grok 4.7&lt;/td&gt;&lt;td&gt;$0.30&lt;/td&gt;&lt;td&gt;140k in / 3k out, under the 200k threshold&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://docs.x.ai/developers/pricing&quot;&gt;xAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, map-reduce&lt;/td&gt;&lt;td&gt;gpt-6.1-sol&lt;/td&gt;&lt;td&gt;$0.43&lt;/td&gt;&lt;td&gt;14 chunk calls + 1 reduce call; 149.8k in / 12.8k out&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, map-reduce&lt;/td&gt;&lt;td&gt;gpt-6-luna&lt;/td&gt;&lt;td&gt;$0.021&lt;/td&gt;&lt;td&gt;14 chunk calls + 1 reduce call&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, map-reduce&lt;/td&gt;&lt;td&gt;Claude Sonnet 5.5&lt;/td&gt;&lt;td&gt;$0.43&lt;/td&gt;&lt;td&gt;14 chunk calls + 1 reduce call&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, map-reduce&lt;/td&gt;&lt;td&gt;Claude Haiku 5.5&lt;/td&gt;&lt;td&gt;$0.11&lt;/td&gt;&lt;td&gt;15 calls, each prompt under 100k&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, map-reduce&lt;/td&gt;&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;&lt;td&gt;$0.16&lt;/td&gt;&lt;td&gt;14 chunk calls + 1 reduce call&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;Google&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, map-reduce&lt;/td&gt;&lt;td&gt;DeepSeek V4.1-Flash&lt;/td&gt;&lt;td&gt;$0.030&lt;/td&gt;&lt;td&gt;off-peak window, 15 calls&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://api-docs.deepseek.com/quick_start/pricing&quot;&gt;DeepSeek&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 2, map-reduce&lt;/td&gt;&lt;td&gt;Grok 4.7&lt;/td&gt;&lt;td&gt;$0.38&lt;/td&gt;&lt;td&gt;14 chunk calls + 1 reduce call&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://docs.x.ai/developers/pricing&quot;&gt;xAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 3, agent loop&lt;/td&gt;&lt;td&gt;gpt-6.1-sol&lt;/td&gt;&lt;td&gt;$0.25&lt;/td&gt;&lt;td&gt;8 turns; 230k in (192.5k read, 37.5k written) / 6k out&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 3, agent loop&lt;/td&gt;&lt;td&gt;gpt-6-luna&lt;/td&gt;&lt;td&gt;$0.013&lt;/td&gt;&lt;td&gt;8 turns; 230k in (192.5k read, 37.5k written) / 6k out&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 3, agent loop&lt;/td&gt;&lt;td&gt;Claude Sonnet 5.5&lt;/td&gt;&lt;td&gt;$0.25&lt;/td&gt;&lt;td&gt;8 turns; 192.5k reads and 37.5k 5-minute writes / 6k out&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 3, agent loop&lt;/td&gt;&lt;td&gt;Claude Haiku 5.5&lt;/td&gt;&lt;td&gt;$0.013&lt;/td&gt;&lt;td&gt;8 turns, every prompt under 100k&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 3, agent loop&lt;/td&gt;&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;&lt;td&gt;$0.065&lt;/td&gt;&lt;td&gt;8 turns; 230k in (192.5k cached) / 6k out, storage excluded&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;Google&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 3, agent loop&lt;/td&gt;&lt;td&gt;DeepSeek V4.1-Flash&lt;/td&gt;&lt;td&gt;$0.0098&lt;/td&gt;&lt;td&gt;off-peak window; 8 turns&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://api-docs.deepseek.com/quick_start/pricing&quot;&gt;DeepSeek&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Job 3, agent loop&lt;/td&gt;&lt;td&gt;Grok 4.7&lt;/td&gt;&lt;td&gt;$0.21&lt;/td&gt;&lt;td&gt;8 turns; 230k in (192.5k cached) / 6k out&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://docs.x.ai/developers/pricing&quot;&gt;xAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;Two pairs of rows are identical on purpose. gpt-6.1-sol and Claude Sonnet 5.5 both list $2.00 per million input tokens, $10.00 output, $0.10 cached input and a $2.50 cache write (&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI pricing&lt;/a&gt; and &lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic pricing&lt;/a&gt;, read 2026-10-09); gpt-6-luna and Claude Haiku 5.5 both list $0.10, $0.50, $0.01 and $0.125.&lt;/p&gt;
&lt;h2 id=&quot;the-same-jobs-with-each-vendors-biggest-discount&quot;&gt;The same jobs with each vendor’s biggest discount&lt;/h2&gt;
&lt;p&gt;Batch is the biggest published discount: 50% on OpenAI (&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;pricing page&lt;/a&gt;, read 2026-10-09), on Anthropic (“a 50% discount on both input and output tokens”, &lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;pricing page&lt;/a&gt;, read 2026-10-09) and on Google’s paid tier (“Batch API (50% cost reduction)”, &lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;pricing page&lt;/a&gt;, read 2026-10-09), where the Batch column halves the cached-input rate too. Not everywhere: xAI’s batch section gives 20% to grok-4.3 and the grok-4.20 variants and says unlisted models have no batch discount, which leaves Grok 4.7 at full price (&lt;a href=&quot;https://docs.x.ai/developers/pricing&quot;&gt;xAI pricing&lt;/a&gt;, read 2026-10-09). DeepSeek publishes no batch tier.&lt;/p&gt;
&lt;div class=&quot;table-wrap&quot;&gt;




































































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Job 1, cached&lt;/th&gt;&lt;th&gt;Job 2, one shot&lt;/th&gt;&lt;th&gt;Job 3, loop&lt;/th&gt;&lt;th&gt;Batch discount&lt;/th&gt;&lt;th&gt;Price source (read 2026-10-09)&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;gpt-6.1-sol&lt;/td&gt;&lt;td&gt;$0.22&lt;/td&gt;&lt;td&gt;$0.16&lt;/td&gt;&lt;td&gt;$0.12&lt;/td&gt;&lt;td&gt;50%&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;gpt-6-luna&lt;/td&gt;&lt;td&gt;$0.012&lt;/td&gt;&lt;td&gt;$0.0078&lt;/td&gt;&lt;td&gt;$0.0067&lt;/td&gt;&lt;td&gt;50%&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Claude Sonnet 5.5&lt;/td&gt;&lt;td&gt;$0.22&lt;/td&gt;&lt;td&gt;$0.16&lt;/td&gt;&lt;td&gt;$0.12&lt;/td&gt;&lt;td&gt;50%&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Claude Haiku 5.5&lt;/td&gt;&lt;td&gt;$0.012&lt;/td&gt;&lt;td&gt;$0.039&lt;/td&gt;&lt;td&gt;$0.0067&lt;/td&gt;&lt;td&gt;50%&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;&lt;td&gt;$0.088&lt;/td&gt;&lt;td&gt;$0.058&lt;/td&gt;&lt;td&gt;$0.033&lt;/td&gt;&lt;td&gt;50%&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;Google&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;DeepSeek V4.1-Flash&lt;/td&gt;&lt;td&gt;$0.028&lt;/td&gt;&lt;td&gt;$0.023&lt;/td&gt;&lt;td&gt;$0.0098&lt;/td&gt;&lt;td&gt;none published&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://api-docs.deepseek.com/quick_start/pricing&quot;&gt;DeepSeek&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Grok 4.7&lt;/td&gt;&lt;td&gt;$0.53&lt;/td&gt;&lt;td&gt;$0.30&lt;/td&gt;&lt;td&gt;$0.21&lt;/td&gt;&lt;td&gt;none&lt;/td&gt;&lt;td&gt;&lt;a href=&quot;https://docs.x.ai/developers/pricing&quot;&gt;xAI&lt;/a&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;Claude Haiku 5.5’s one-shot cell is the outlier: batch brings it to $0.039, still three times its cached Job 1 of $0.012, because a 140,000-token prompt sits above the 100,000-token step (&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic pricing&lt;/a&gt;, read 2026-10-09).&lt;/p&gt;
&lt;h2 id=&quot;what-the-numbers-say&quot;&gt;What the numbers say&lt;/h2&gt;
&lt;p&gt;Every figure here is arithmetic on the rows above, whose rates are linked (&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI&lt;/a&gt;, &lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic&lt;/a&gt;, &lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;Google&lt;/a&gt;, &lt;a href=&quot;https://api-docs.deepseek.com/quick_start/pricing&quot;&gt;DeepSeek&lt;/a&gt;, &lt;a href=&quot;https://docs.x.ai/developers/pricing&quot;&gt;xAI&lt;/a&gt;, read 2026-10-09).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Output tokens decide the bill at the frontier tier.&lt;/strong&gt; In cached Job 1 on gpt-6.1-sol, $0.15 of the $0.43 is output (35%) and $0.04 is cached instruction tokens (9%).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The spread inside one vendor is bigger than the spread between vendors.&lt;/strong&gt; Cached Job 1 is $0.024 on gpt-6-luna and Haiku 5.5, $0.43 on gpt-6.1-sol and Sonnet 5.5, $0.53 on Grok 4.7: 22x from dearest to cheapest, 18x between the cheap pair and the frontier pair. At the same tier, the vendor changes nothing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The agent loop is the cheapest job here.&lt;/strong&gt; Job 3 is $0.25 on gpt-6.1-sol against $0.31 for the one-shot summary and $1.19 for Job 1. On the small models the order flips: Haiku 5.5 pays $0.078 for the one-shot summary against $0.013 for the loop, because 140,000 tokens crosses its threshold and 230,000 over eight turns does not.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Map-reduce costs more than one shot.&lt;/strong&gt; On the two $2/$10 models it is $0.43 against $0.31, and the gap is all output tokens, 12,800 against 3,000. Chunking buys reliability and pays for it in output.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The absolute numbers are small.&lt;/strong&gt; The main table spans $0.0098 to $1.19; batch pulls the floor to $0.0067. At this scale, a cost problem is usually a volume problem.&lt;/p&gt;
&lt;h2 id=&quot;the-variables-that-break-the-conversion&quot;&gt;The variables that break the conversion&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Thresholds re-price the whole request.&lt;/strong&gt; Claude Haiku 5.5 charges $0.10 and $0.50 per million input and output tokens up to 100,000 prompt tokens, and $0.50 and $2.50 above it; the length counts cache reads and writes, and each request is priced on its own (&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic pricing&lt;/a&gt;, read 2026-10-09). Job 2’s 140,000-token prompt pays the higher rate on everything: $0.078, against $0.016 under the cap. xAI applies long-context rates to all tokens in a request once the prompt passes 200,000, where Grok 4.7 costs $4.00 input, $1.00 cached input and $12.00 output instead of $2.00, $0.50 and $6.00 (&lt;a href=&quot;https://docs.x.ai/developers/pricing&quot;&gt;xAI pricing&lt;/a&gt;, read 2026-10-09). Gemini 3.1 Pro Preview steps from $2.00 to $4.00 input per million above 200,000 tokens (&lt;a href=&quot;https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing&quot;&gt;Gemini Enterprise Agent Platform pricing&lt;/a&gt;, read 2026-10-09). OpenAI’s line is 272,000 input tokens, published in a column tooltip; above it gpt-6.1-sol’s input doubles to $4.00 and its output to $15.00 (&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI pricing&lt;/a&gt;, read 2026-10-09). All three jobs stay under every line; Job 2 is halfway to OpenAI’s.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tokenizer differences.&lt;/strong&gt; Anthropic notes that Claude 4.7 and later use a tokenizer that “produces approximately 30% more tokens for the same text” (&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic pricing&lt;/a&gt;, read 2026-10-09). A comparison built on one words-per-token ratio, as mine is, is off by that much for Anthropic.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Time of day.&lt;/strong&gt; DeepSeek’s peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday, excluding Chinese public holidays; every other hour is off-peak at half price (&lt;a href=&quot;https://api-docs.deepseek.com/quick_start/pricing&quot;&gt;DeepSeek Models &amp;#x26; Pricing&lt;/a&gt;, read 2026-10-09). Job 1 uncached is $0.087 off-peak and $0.174 at peak. Off-peak is 133 of the week’s 168 hours.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cache economics differ in shape as well as in rate.&lt;/strong&gt; Anthropic’s cache hit is 10% of the input price on most models, 5% on Opus 5.5 and Sonnet 5.5, 2.5% on Fable 5.1 and Mythos 5.1; writes cost 1.25x for five minutes and 2x for an hour, so a 5-minute cache pays for itself after one read (&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic pricing&lt;/a&gt;, read 2026-10-09). DeepSeek’s hit price is $0.003 per million against $0.15 for a miss (&lt;a href=&quot;https://api-docs.deepseek.com/quick_start/pricing&quot;&gt;DeepSeek pricing&lt;/a&gt;, read 2026-10-09). Google charges no write fee but charges storage, $0.50 per million tokens per hour on Gemini 3.8 Flash through 2026-12-31 and $4.50 on Gemini 3.1 Pro (&lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;Gemini pricing&lt;/a&gt; and &lt;a href=&quot;https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing&quot;&gt;Gemini Enterprise Agent Platform pricing&lt;/a&gt;, read 2026-10-09); a cached 140,000-token document runs $0.07 an hour on Flash and $0.63 on Pro.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Platform fees sit outside the token meter.&lt;/strong&gt; OpenAI bills $2.50 per 1,000 tool calls and $10 per 1,000 web searches, so Job 3’s eight calls cost $0.02, and Anthropic’s tool-use prompt adds 286 to 675 tokens (&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI pricing&lt;/a&gt; and &lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic pricing&lt;/a&gt;, read 2026-10-09). xAI bills X search per item fetched: $5 per 1,000 posts, $10 per 1,000 profiles (&lt;a href=&quot;https://docs.x.ai/developers/pricing&quot;&gt;xAI pricing&lt;/a&gt;, read 2026-10-09).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Latency and geography carry multipliers.&lt;/strong&gt; OpenAI’s Fast mode, renamed from Priority on 2026-07-30, is 2x, and data-residency endpoints add 10%; Gemini 3.8 Flash’s Priority tier is $1.35 input against $0.75 standard (&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI pricing&lt;/a&gt; and &lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;Gemini pricing&lt;/a&gt;, read 2026-10-09). xAI’s priority tier is 2x and its US regional endpoint 1.1x; Anthropic charges 1.1x for US-only inference; regional Bedrock and Google Cloud endpoints carry a 10% premium (&lt;a href=&quot;https://docs.x.ai/developers/pricing&quot;&gt;xAI pricing&lt;/a&gt; and &lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic pricing&lt;/a&gt;, read 2026-10-09).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The list price itself moves.&lt;/strong&gt; Google quotes its Flash prices “through December 31, 2026” and doubles them on 2027-01-01: $0.75 to $1.50 input, $3.75 to $7.50 output, $0.075 to $0.15 cached input, $0.50 to $1.00 per million tokens per hour of cache storage (&lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;Gemini pricing&lt;/a&gt;, read 2026-10-09). Anthropic’s $2/$10 for Claude Sonnet 5 was launched as introductory pricing through 2026-08-31 and is now standard, with the rise to $3/$15 cancelled (&lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/pricing&quot;&gt;Anthropic pricing&lt;/a&gt;, read 2026-10-09). OpenAI’s promotional pricing for GPT-5.6 Sol runs at least through 2026-11-21, on a model not in this table (&lt;a href=&quot;https://developers.openai.com/api/docs/pricing&quot;&gt;OpenAI pricing&lt;/a&gt;, read 2026-10-09). Any per-task table is a snapshot with an expiry date.&lt;/p&gt;
&lt;h2 id=&quot;what-i-could-not-check&quot;&gt;What I could not check&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;I ran none of these three jobs.&lt;/strong&gt; The token counts are stated assumptions, and a measured run would move every number here.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cache-write pricing for Google, DeepSeek and xAI.&lt;/strong&gt; I found no published write price for any of the three, so I charged zero. xAI’s prompt-caching page lists a billing rate per token type with no write line; Google’s context-caching price is a read rate plus hourly storage; DeepSeek’s table has only a hit price and a miss price. If any of them does charge for writes, its cached rows are too low.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Whether the Haiku 5.5 threshold behaves as written at the boundary.&lt;/strong&gt; The page counts cache reads and writes in a prompt’s length, which would mean a cached 140,000-token document pays the higher rate. I did not test it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Real token counts.&lt;/strong&gt; I used Google’s words-per-token ratio across every vendor; Anthropic’s own note about a 30% tokenizer difference is a warning that this is approximate.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Volume, committed-use and negotiated enterprise pricing, free tiers and trial credits&lt;/strong&gt; are outside this table. The price on the page is the list price.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Resale through cloud platforms.&lt;/strong&gt; Anthropic bills Bedrock and Google Cloud in consumption units, and Gemini Enterprise Agent Platform prices are not the Gemini API prices: on 2026-10-09 the two pages disagreed on cache storage, $0.50 per million tokens per hour (&lt;a href=&quot;https://ai.google.dev/gemini-api/docs/pricing&quot;&gt;Gemini pricing&lt;/a&gt;) against $1.00 for one model (&lt;a href=&quot;https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing&quot;&gt;Gemini Enterprise Agent Platform pricing&lt;/a&gt;). The route you buy through can change the number.&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Why there is no Western mini program</title><link>https://notebookfield.com/posts/why-there-is-no-western-mini-program/</link><guid isPermaLink="true">https://notebookfield.com/posts/why-there-is-no-western-mini-program/</guid><description>A mini program inherits its identity and its payment rail from the app hosting it. That mechanism, and what the four Western equivalents do instead.</description><pubDate>Fri, 09 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The usual version of this question is about super apps, and it has been answered at length by people with better data than mine; mine is narrower. A mini program is a small bundle of code that runs inside somebody else’s application, and the question is which parts of that arrangement exist in what Apple and Google ship.&lt;/p&gt;
&lt;p&gt;I read the vendors’ own documentation for WeChat Mini Programs, Apple App Clips, Google Play Instant and Progressive Web Apps, and compared them on five dimensions: install and size, identity, payment, entry point, distribution and review.&lt;/p&gt;
&lt;p&gt;One mechanism carries most of the answer. A mini program brings neither its own user account nor its own checkout. It borrows both from the application hosting it.&lt;/p&gt;
&lt;h2 id=&quot;the-five-dimensions-side-by-side&quot;&gt;The five dimensions, side by side&lt;/h2&gt;
&lt;div class=&quot;table-wrap&quot;&gt;














































&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;/th&gt;&lt;th&gt;WeChat Mini Program&lt;/th&gt;&lt;th&gt;Progressive web app&lt;/th&gt;&lt;th&gt;Native app (iOS/Android)&lt;/th&gt;&lt;th&gt;App Clip (iOS)&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Install and size&lt;/td&gt;&lt;td&gt;Runs inside WeChat. Main package 2 MB maximum, all packages together 30 MB (20 MB for programs developed by a third-party provider) (&lt;a href=&quot;https://developers.weixin.qq.com/miniprogram/dev/framework/subpackages.html&quot;&gt;WeChat subpackages&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;td&gt;No package at all. Chrome offers install after HTTPS plus a manifest carrying a name, 192 px and 512 px icons, &lt;code&gt;start_url&lt;/code&gt; and a display mode (&lt;a href=&quot;https://web.dev/articles/install-criteria&quot;&gt;web.dev install criteria&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;td&gt;Full binary from the store. Apple states the App Store limits the size of apps installable over a mobile connection and gives no number on that page (&lt;a href=&quot;https://developer.apple.com/documentation/xcode/reducing-your-app-s-size&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09); Google states a compressed download limit of 200 MB for apps published as app bundles (&lt;a href=&quot;https://developer.android.com/topic/performance/reduce-apk-size&quot;&gt;Google&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;td&gt;10 MB uncompressed on iOS 15 and earlier, 15 MB on iOS 16 and earlier, 100 MB on iOS 17 or later under conditions (&lt;a href=&quot;https://developer.apple.com/documentation/appclip/choosing-the-right-functionality-for-your-app-clip&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Identity&lt;/td&gt;&lt;td&gt;&lt;code&gt;openid&lt;/code&gt; per mini program, &lt;code&gt;unionid&lt;/code&gt; shared across every app under one Open Platform account. Both come from &lt;code&gt;wx.login&lt;/code&gt; plus &lt;code&gt;code2Session&lt;/code&gt;, which WeChat documents as returning the UnionID “without user authorization” (&lt;a href=&quot;https://developers.weixin.qq.com/miniprogram/en/dev/framework/open-ability/login.html&quot;&gt;login&lt;/a&gt;, &lt;a href=&quot;https://developers.weixin.qq.com/miniprogram/en/dev/framework/open-ability/union-id.html&quot;&gt;UnionID&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;td&gt;Each origin keeps its own account. The browser hands nothing to a site silently&lt;/td&gt;&lt;td&gt;Account per app. Sign in with Apple where the app offers it, as an explicit step by the user (&lt;a href=&quot;https://developer.apple.com/sign-in-with-apple&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;td&gt;Sign in with Apple, the keychain and CloudKit, shared with the parent app (&lt;a href=&quot;https://developer.apple.com/documentation/appclip/sharing-data-between-your-app-clip-and-your-full-app&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Payment&lt;/td&gt;&lt;td&gt;WeChat Pay. A merchant number is bound to the mini program’s AppID, at most 50 AppIDs per merchant number (&lt;a href=&quot;https://pay.weixin.qq.com/doc/v3/merchant/4013287504&quot;&gt;WeChat Pay&lt;/a&gt;, page updated 2025-07-02). Virtual goods must go through the mini program virtual-payment channel, where “the platform charges a technical service fee” (&lt;a href=&quot;https://developers.weixin.qq.com/miniprogram/dev/platform-capabilities/en/business-capabilities/virtual-payment.html&quot;&gt;WeChat&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;td&gt;Card and wallet APIs called by the merchant, on the merchant’s own acquiring contract&lt;/td&gt;&lt;td&gt;In-app purchase is required to unlock features. Apple’s guidelines name QR codes among the mechanisms that may not be used instead (&lt;a href=&quot;https://developer.apple.com/app-store/review/guidelines/&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;td&gt;Apple’s in-app purchase. Hosts in the Mini Apps Partner Program “earn 85% of qualifying In-App Purchase sales within qualifying mini apps” (&lt;a href=&quot;https://developer.apple.com/programs/mini-apps-partner/&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Entry&lt;/td&gt;&lt;td&gt;Documented entry paths carry IDs: 1011 scan a QR code, 1012 long-press an image, 1013 pick a code from the album, 1023 the Android home-screen icon, 1005 and 1006 search boxes, 1026 the nearby list, 1183 search inside the PC WeChat mini program panel (&lt;a href=&quot;https://developers.weixin.qq.com/miniprogram/dev/reference/scene-list.html&quot;&gt;WeChat scene values&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;td&gt;A URL, a search result, or a home-screen shortcut the user adds by hand. On iOS there is no install prompt at all (&lt;a href=&quot;https://web.dev/learn/pwa/installation/&quot;&gt;web.dev&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;td&gt;Store search and the home-screen icon&lt;/td&gt;&lt;td&gt;An App Clip Code, NFC tag or QR code at a physical location; Siri suggestions; Maps; Messages; a Smart App Banner on a website in Safari (&lt;a href=&quot;https://developer.apple.com/documentation/appclip/configuring-the-launch-experience-of-your-app-clip&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Distribution and review&lt;/td&gt;&lt;td&gt;Published from the mini program console. Virtual payment requires a certified entity: an enterprise, institution or individual merchant (&lt;a href=&quot;https://developers.weixin.qq.com/miniprogram/dev/platform-capabilities/en/business-capabilities/virtual-payment.html&quot;&gt;WeChat&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;td&gt;Nothing is submitted. It is HTTPS plus a manifest&lt;/td&gt;&lt;td&gt;Store review against the App Review Guidelines&lt;/td&gt;&lt;td&gt;No listing of its own. It ships inside a full app, which the App Clip documentation says “must include the same functionality as the App Clip” (&lt;a href=&quot;https://developer.apple.com/documentation/appclip/choosing-the-right-functionality-for-your-app-clip&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09); Guideline 2.5.16(a) adds that all App Clip features must be in the main app binary and that “App Clips cannot contain advertising” (&lt;a href=&quot;https://developer.apple.com/app-store/review/guidelines/&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The Android equivalent is gone. “Starting December 2025, Instant Apps cannot be published through Google Play, and all Google Play services Instant APIs will no longer work” (&lt;a href=&quot;https://developer.android.com/topic/google-play-instant&quot;&gt;Google Play Instant&lt;/a&gt;, last updated 2026-06-24).&lt;/p&gt;
&lt;h2 id=&quot;identity-arrives-before-the-first-screen&quot;&gt;Identity arrives before the first screen&lt;/h2&gt;
&lt;p&gt;The WeChat login flow is one call on the client and one on the server. &lt;code&gt;wx.login&lt;/code&gt; returns a code, the server exchanges it through &lt;code&gt;code2Session&lt;/code&gt;, and the reply carries an &lt;code&gt;openid&lt;/code&gt; for that user in that mini program, a &lt;code&gt;unionid&lt;/code&gt; and a session key (&lt;a href=&quot;https://developers.weixin.qq.com/miniprogram/en/dev/framework/open-ability/login.html&quot;&gt;WeChat&lt;/a&gt;, read 2026-10-09).&lt;/p&gt;
&lt;p&gt;The UnionID page says what this costs the user: the developer gets it “without user authorization” (&lt;a href=&quot;https://developers.weixin.qq.com/miniprogram/en/dev/framework/open-ability/union-id.html&quot;&gt;WeChat&lt;/a&gt;, read 2026-10-09). No phone number, no password, no consent sheet. The &lt;code&gt;openid&lt;/code&gt; is scoped to one mini program; the &lt;code&gt;unionid&lt;/code&gt; is identical across every mobile app, website app, official account and mini program bound to the same Open Platform account, so a merchant with several surfaces sees one person.&lt;/p&gt;
&lt;p&gt;Compare the alternatives. On the web, an origin cannot read an identifier out of the browser; it asks the user to sign in and then owns a credential to protect. On iOS, an App Clip can use Sign in with Apple and the keychain, which Apple’s documentation points developers at sharing with the full app, but those are technologies an app opts into rather than an identifier handed over at launch.&lt;/p&gt;
&lt;p&gt;That is why the first screen of a mini program is the product. Onboarding was paid for at install time, by the host.&lt;/p&gt;
&lt;h2 id=&quot;payment-is-a-binding-between-two-accounts&quot;&gt;Payment is a binding between two accounts&lt;/h2&gt;
&lt;p&gt;Paying inside a mini program does not mean the mini program holds money. WeChat Pay binds a merchant number to the mini program’s AppID. One merchant number can bind at most 50 AppIDs, the binding starts on the merchant side and is confirmed from the mini program side, and once established cannot be unbound (&lt;a href=&quot;https://pay.weixin.qq.com/doc/v3/merchant/4013287504&quot;&gt;WeChat Pay&lt;/a&gt;, page updated 2025-07-02). The rail and the code package are joined by an account relationship neither side can create alone.&lt;/p&gt;
&lt;p&gt;Virtual goods get a second layer. The documentation is blunt: virtual currency, unlocked features, subscription content, paid services, tips and virtual gifts “require integration with Mini Program’s virtual payment system for both purchase and payment”, and the platform charges a technical service fee on the amount (&lt;a href=&quot;https://developers.weixin.qq.com/miniprogram/dev/platform-capabilities/en/business-capabilities/virtual-payment.html&quot;&gt;WeChat&lt;/a&gt;, read 2026-10-09). Physical goods go through ordinary WeChat Pay; digital ones go through the host’s channel, with the host’s cut.&lt;/p&gt;
&lt;p&gt;Apple’s version of that cut is recent and documented. The Mini Apps Partner Program was announced on 13 November 2025 (&lt;a href=&quot;https://developer.apple.com/news/?id=xcz1s7cz&quot;&gt;Apple&lt;/a&gt;, 2025-11-13): host apps carrying mini apps can join, earn 85 percent of qualifying in-app purchase sales inside those mini apps, and must use Apple’s in-app purchase system for them (&lt;a href=&quot;https://developer.apple.com/programs/mini-apps-partner/&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09). Admission is not automatic: a host needs the Advanced Commerce API and a manifest describing itself and its mini apps.&lt;/p&gt;
&lt;p&gt;Apple’s App Review Guidelines still require in-app purchase to unlock features or functionality, and the same sentence lists QR codes among the mechanisms that cannot substitute for it (&lt;a href=&quot;https://developer.apple.com/app-store/review/guidelines/&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09). What changed, on 1 May 2025, is that Guidelines 3.1.1, 3.1.1(a), 3.1.3 and 3.1.3(a) were updated for the United States storefront after a court decision about buttons and external links (&lt;a href=&quot;https://developer.apple.com/news/?id=9txfddzf&quot;&gt;Apple&lt;/a&gt;, 2025-05-01).&lt;/p&gt;
&lt;h2 id=&quot;the-qr-code-is-the-offline-switch&quot;&gt;The QR code is the offline switch&lt;/h2&gt;
&lt;p&gt;A WeChat developer does not have to choose between a printed code and a link shared in a chat. The official scene-value table lists both, with IDs: 1011 scanning a QR code, 1012 long-pressing an image, 1013 picking one from the album, 1023 the Android home-screen icon, 1005 and 1006 the two search boxes, 1026 the nearby list, 1183 search inside the PC WeChat mini program panel (&lt;a href=&quot;https://developers.weixin.qq.com/miniprogram/dev/reference/scene-list.html&quot;&gt;WeChat&lt;/a&gt;, read 2026-10-09).&lt;/p&gt;
&lt;p&gt;Two details make the printed code cheap. Codes come from an API, and WeChat states every generated mini program code is permanently valid. The “one item, one code” interface has no quantity limit and needs a smaller printed area than an ordinary URL code (&lt;a href=&quot;https://developers.weixin.qq.com/miniprogram/dev/framework/open-ability/qr-code.html&quot;&gt;WeChat&lt;/a&gt;, read 2026-10-09). A code already printed and pointing at a URL can be configured to open a chosen page inside a mini program (&lt;a href=&quot;https://developers.weixin.qq.com/miniprogram/introduction/qrcode.html&quot;&gt;WeChat&lt;/a&gt;, read 2026-10-09). The person scans what is already on the table.&lt;/p&gt;
&lt;p&gt;App Clips have the same physical invocations: an App Clip Code, an NFC tag, a QR code at a physical location (&lt;a href=&quot;https://developer.apple.com/documentation/appclip/configuring-the-launch-experience-of-your-app-clip&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09). Two constraints pull the other way.&lt;/p&gt;
&lt;p&gt;The first constraint trades size against physical entry. The 100 MB ceiling for iOS 17 or later carries four conditions: the App Clip “only supports digital invocations” and not “physical invocations such as App Clip Codes, QR codes, or NFC tags” (&lt;a href=&quot;https://developer.apple.com/documentation/appclip/choosing-the-right-functionality-for-your-app-clip&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09), it is used where a reliable connection is likely, and it “doesn’t support iOS 16 and earlier”. The exception Apple names is a demo link from App Store Connect, which may use the 100 MB limit and still support App Clip Codes, NFC tags and QR codes. An App Clip with a counter code stays inside the smaller limit.&lt;/p&gt;
&lt;p&gt;The second concerns persistence. Apple’s Human Interface Guidelines say App Clips “remain on the device for a limited amount of time”, and that only App Clip Codes produced in App Store Connect or with Apple’s own code generator are approved for use (&lt;a href=&quot;https://developer.apple.com/design/human-interface-guidelines/app-clips&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09). A merchant cannot print its own.&lt;/p&gt;
&lt;h2 id=&quot;who-reviews-it-and-who-is-answerable&quot;&gt;Who reviews it, and who is answerable&lt;/h2&gt;
&lt;p&gt;Reviews differ in kind as well as in strictness. A PWA is submitted to nobody; Chrome installs it once its manifest and HTTPS satisfy the criteria (&lt;a href=&quot;https://web.dev/articles/install-criteria&quot;&gt;web.dev&lt;/a&gt;, read 2026-10-09). An App Clip cannot be submitted alone. It ships inside a full app that must already contain the same functionality, cannot carry advertising, and is reviewed as part of that app (&lt;a href=&quot;https://developer.apple.com/app-store/review/guidelines/&quot;&gt;Apple&lt;/a&gt;, read 2026-10-09). A mini program is submitted inside WeChat, and the entity requirement arrives with payment rather than with the code: to sell virtual items, the program must be a certified enterprise, institution or individual merchant (&lt;a href=&quot;https://developers.weixin.qq.com/miniprogram/dev/platform-capabilities/en/business-capabilities/virtual-payment.html&quot;&gt;WeChat&lt;/a&gt;, read 2026-10-09).&lt;/p&gt;
&lt;p&gt;The commercial consequence: in the App Store, a mini-app-like product reaches users through a host that already has an account relationship and a reviewed app behind it. In WeChat, the party needing the corporate identity, the certification and the merchant binding is the mini program author.&lt;/p&gt;
&lt;h2 id=&quot;where-the-four-pieces-sit&quot;&gt;Where the four pieces sit&lt;/h2&gt;
&lt;p&gt;In WeChat, four things a mini program depends on sit with one owner: the identity system, the payment rail, the entry surface (camera, search, chat, home screen, desktop) and the distribution channel with its review. An author plugs into a host that has all four and pays for the parts it uses.&lt;/p&gt;
&lt;p&gt;In the Apple and Google stacks those four are held by parties with different interests. The operating system vendor issues the platform account and runs the payment system; the browser owns the entry surface for web content; the stores own distribution and review; the rail for physical goods sits with the merchant’s bank relationship.&lt;/p&gt;
&lt;p&gt;No single party there can hand all four to a third party. Apple’s answer in November 2025 was to supply the missing pieces itself and keep the commission: hosts get Apple’s in-app purchase for mini apps and keep 85 percent of qualifying sales, in exchange for approval, a manifest and Apple’s review.&lt;/p&gt;
&lt;p&gt;The mechanism that makes mini programs work is therefore an ownership arrangement rather than a technology gap. Reproducing it outside WeChat means finding one party that holds the users, the identity, the payment rail and the entry surface, and getting it to open all four to third-party code at a price the third party will pay. In the West, the two parties that hold those pieces chose to be the host themselves.&lt;/p&gt;
&lt;h2 id=&quot;how-i-checked-this&quot;&gt;How I checked this&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Every figure and quote comes from documentation published by the vendor it describes: WeChat’s mini program and WeChat Pay docs, Apple’s developer documentation, guidelines and news items, Google’s Android documentation, and web.dev.&lt;/li&gt;
&lt;li&gt;All pages were fetched on 2026-10-09. Most carry no publication date, so a date beside a source is my reading date unless the page states its own: the WeChat Pay AppID binding page (updated 2025-07-02), the Google Play Instant page (last updated 2026-06-24), Apple’s updated-guidelines news item (2025-05-01), Apple’s Mini Apps Partner Program news item (2025-11-13) and Tencent’s write-up of the 2022 WeChat Open Class (2022-01-07).&lt;/li&gt;
&lt;li&gt;Size limits and scene-value IDs come from the pages that state them: WeChat’s subpackage page for 2 MB and 30 MB, Apple’s App Clip page for the 10/15/100 MB table, WeChat’s scene-value table for the ID numbers.&lt;/li&gt;
&lt;li&gt;Google Play Instant’s page would not load in a plain HTTP client, so I read it through the Internet Archive and then through a rendering fetch that returned the live page.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&quot;what-i-could-not-check&quot;&gt;What I could not check&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The fee WeChat charges on virtual payments.&lt;/strong&gt; The page says a fee is charged on the payment amount and gives no percentage.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The standard WeChat Pay rate for physical goods.&lt;/strong&gt; I found no official rate page; what I reached describes partner fee schemes and refers to a merchant rate without the base number.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The extra steps for virtual payment on iOS.&lt;/strong&gt; WeChat’s virtual payment page says iOS “requires additional activation and adaptation procedures” and points at separate documentation; I could not load that page, so I do not know whether an Apple-side charge is involved.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Whether an App Clip can take payment for physical goods through Apple Pay.&lt;/strong&gt; Not stated on the App Clip pages I read.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Any current count of mini programs or mini-program daily users.&lt;/strong&gt; Tencent’s release discloses Weixin monthly accounts and names Mini Games and Mini Shops, but carries no mini-program count. The most recent first-party figure I found: more than 450 million daily users at the end of 2021, from Tencent’s write-up of the 2022 WeChat Open Class (&lt;a href=&quot;https://www.tencent.com/zh-cn/articles/2201267.html&quot;&gt;Tencent&lt;/a&gt;, 2022-01-07). That figure is four years old.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What “certified Mini Program” requires in practice&lt;/strong&gt;, beyond the entity types named on the virtual payment page, and how long WeChat’s review takes.&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item><item><title>Twenty hot lists, no API keys</title><link>https://notebookfield.com/posts/twenty-hot-lists-no-api-keys/</link><guid isPermaLink="true">https://notebookfield.com/posts/twenty-hot-lists-no-api-keys/</guid><description>An aggregator that collects 20 trending lists every night. Six sources broke in ways worth keeping a record of, and one of them is gone for good.</description><pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Every night at 04:30 a script on this machine collects the current front page of 20 platforms
(Zhihu, Bilibili, V2EX, Hupu, Maoyan, Hacker News, Lobsters, GitHub, Reddit, YouTube, arXiv and a
few tech news feeds), translates the English headlines into Chinese, builds a static site and
pushes it. There is no server and no database. The output is a set of JSON files and a
single-page app that reads them.&lt;/p&gt;
&lt;p&gt;I have no API keys for most of these platforms, and for the ones that do offer keys, I did not
want to create accounts for a hobby project. That constraint produced the most useful part of the
system: a catalogue of what each source actually gives you, and a record of how each one broke.
This is that record. It is more useful than the feature list.&lt;/p&gt;
&lt;h2 id=&quot;the-sources-that-broke-and-what-the-failure-was-really-about&quot;&gt;The sources that broke, and what the failure was really about&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Reddit, first attempt.&lt;/strong&gt; There was a trick where you request
&lt;code&gt;translate.google.com/translate?u=&amp;#x3C;reddit url&gt;&lt;/code&gt; and read the translated page, which passes through
Reddit’s HTML. It worked for months and then started returning a 302 to &lt;code&gt;translate.goog&lt;/code&gt;, which
serves a shell with no posts in it. The failure looked like a parser bug. It was a removed feature.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reddit, second attempt.&lt;/strong&gt; The official Atom feed, &lt;code&gt;r/popular/hot/.rss&lt;/code&gt;, works and needs no key.
Two traps: &lt;code&gt;geo_filter=GLOBAL&lt;/code&gt; has to be present, or the feed comes back localised to the country
your request exits from, and mine exits from a node that made the whole page German. And the feed
carries no score and no comment count, so those fields are simply absent from my data rather than
zero.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Reddit, third attempt, which is where it ends.&lt;/strong&gt; Scores are available through the JSON API if
you have OAuth credentials. I opened the app-creation form, submitted it, and got
&lt;code&gt;{&quot;success&quot;: true}&lt;/code&gt; back, while the applications list stayed empty. Their documentation has since
moved new legacy-Data-API apps to an application process for moderation use cases. Unauthenticated
&lt;code&gt;.json&lt;/code&gt; returns 403 from anything that looks like a datacenter address: direct, through a proxy,
and with a browser TLS fingerprint that makes the request look like Chrome. Two public mirrors
were dead ends as well, one of which now wants to be paid and explicitly refuses automated
clients, and the other of which returns &lt;code&gt;/r/popular&lt;/code&gt; posts from 2015.&lt;/p&gt;
&lt;p&gt;There is a path that works: send the request with a logged-in browser’s cookies. I measured it,
got a real score back, and decided not to use it. The data is not worth building the pipeline on
something the platform’s terms do not allow.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;arXiv.&lt;/strong&gt; The API started answering with a 14-byte body, &lt;code&gt;429 &quot;Rate exceeded.&quot;&lt;/code&gt;, for every query,
from every route. Not occasional throttling: the endpoint was closed to my traffic entirely.
The official RSS feeds for cs.AI, cs.LG and cs.CL serve the same listing without a key. The
endpoint is still in the code as a fallback, and I expect to delete it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;36Kr.&lt;/strong&gt; The page is entirely client-rendered, so the extractor’s basic mode returns
&lt;code&gt;{&quot;results&quot;: [], &quot;failed_results&quot;: [{&quot;error&quot;: &quot;Error fetching content&quot;}]}&lt;/code&gt;. The old code read
&lt;code&gt;results[0]&lt;/code&gt; and raised an &lt;code&gt;IndexError&lt;/code&gt;, which the pipeline caught as a platform failure and
handled by keeping yesterday’s data. So for weeks the site showed a platform labelled “AI news”
that was quietly a day or two out of date. The fix was one parameter, &lt;code&gt;extract_depth: advanced&lt;/code&gt;.
The bug was the exception handling that turned a hard failure into a silent stale one.&lt;/p&gt;
&lt;h2 id=&quot;the-night-the-whole-thing-hung&quot;&gt;The night the whole thing hung&lt;/h2&gt;
&lt;p&gt;One night no platform updated. The cron entry reported a timeout after 3600 seconds and the
captured output was empty.&lt;/p&gt;
&lt;p&gt;Nothing was wrong with the network: I could see every platform’s connection being made in the
proxy log at 04:30, and then no connections at all for the next 55 minutes. Nothing had reached
the build step either, which I could tell from one log file’s modification time. The script was
alive and doing nothing, and I had no idea where.&lt;/p&gt;
&lt;p&gt;Two things had made that state undiagnosable. Python buffers stdout when it is not attached to a
terminal, so when the process was killed, every line it had printed died with it. And
&lt;code&gt;urlopen(timeout=N)&lt;/code&gt; applies to each socket read rather than the whole request, so a slow response
can stall for far longer than the timeout suggests.&lt;/p&gt;
&lt;p&gt;The fix is unglamorous: &lt;code&gt;exec &amp;#x3C;/dev/null&lt;/code&gt; so no child can block on stdin, &lt;code&gt;python3 -u&lt;/code&gt; so output
survives a kill, a &lt;code&gt;timeout&lt;/code&gt; around every step, a translation budget, and a step log that records
a timestamp as each stage starts. The next failure will be somewhere specific in that log.&lt;/p&gt;
&lt;h2 id=&quot;a-successful-job-is-not-fresh-data&quot;&gt;A successful job is not fresh data&lt;/h2&gt;
&lt;p&gt;This is the lesson I would keep if I could keep only one. “The task exited 0” and “the data is
from today” are different claims, and a pipeline that only reports the first one will lie to you
quietly for weeks. So there is a separate check: it reads every platform’s JSON, compares its date
with today, and exits non-zero if any source is stale or empty. It does not care that the fetch
succeeded.&lt;/p&gt;
&lt;p&gt;The same principle shows up in the summary step. A model writes the “today’s focus” card, reading
the top five items from each platform, and it is asked to return only item identifiers and a
reason. The script then looks up each identifier and fills in the title and URL from the data it
already has. The model never writes a link, so it cannot invent one. If the summary fails, the old
one stays on the page marked stale.&lt;/p&gt;
&lt;h2 id=&quot;where-the-ordering-is-not-what-it-looks-like&quot;&gt;Where the ordering is not what it looks like&lt;/h2&gt;
&lt;p&gt;Three caveats, since the site presents these lists as rankings:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Hupu’s “hot” list is a weighted stream rather than a sorted list of numbers. There is no score
to sort by.&lt;/li&gt;
&lt;li&gt;36Kr and TechCrunch are editorially ordered. The order is theirs, not a measurement.&lt;/li&gt;
&lt;li&gt;V2EX’s public hot endpoint returns about nine items on a good day.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The site says which list is which on each platform page. A ranking whose method is unstated is
just a list.&lt;/p&gt;
&lt;h2 id=&quot;reproducing-this&quot;&gt;Reproducing this&lt;/h2&gt;
&lt;p&gt;The collector is Python, one function per platform, and the run prints a report. The interesting
part to copy is not the fetching. It is the two checks around it: a freshness check that fails
when data is old even though the job succeeded, and a generation step where the model returns only
identifiers that the script resolves into real URLs.&lt;/p&gt;</content:encoded></item><item><title>Verifying an email signup end to end</title><link>https://notebookfield.com/posts/verifying-an-email-signup/</link><guid isPermaLink="true">https://notebookfield.com/posts/verifying-an-email-signup/</guid><description>The form returned success. The subscriber appeared in the right group. The one thing that failed was in an inbox I do not control.</description><pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;This site has a signup form and a double opt-in. When you submit it, the mailing tool sends you a
confirmation email, and until you click the link in that email you are not on the list.&lt;/p&gt;
&lt;p&gt;I wanted to know whether that actually worked, so I ran the loop once, with a real address, and
checked the state at each step. Here is what “checked” meant, and the failure that only showed up
at the far end.&lt;/p&gt;
&lt;h2 id=&quot;what-the-states-look-like&quot;&gt;What the states look like&lt;/h2&gt;
&lt;p&gt;The tool has two states that matter: &lt;code&gt;Unconfirmed&lt;/code&gt; and &lt;code&gt;Active&lt;/code&gt;. A new signup lands in
&lt;code&gt;Unconfirmed&lt;/code&gt;, does not appear in the default list of subscribers, and does not count towards the
group’s size. That is correct behaviour for double opt-in, and it is also the single most
confusing thing about the setup: the list looks empty right after a successful signup.&lt;/p&gt;
&lt;p&gt;So the check is not “did the count go up”. It is a sequence:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Submit the form on the live site, from the live site’s own page, and read the response. Mine
came back &lt;code&gt;{&quot;success&quot;: true}&lt;/code&gt; and the page replaced the form with a thank-you message.&lt;/li&gt;
&lt;li&gt;Open the subscriber record and read four fields: status &lt;code&gt;Unconfirmed&lt;/code&gt;, the source line naming
the form, the group it was added to, and the activity entry that says it was added.&lt;/li&gt;
&lt;li&gt;Check the campaign report for the confirmation email and confirm its sent counter moved from 1
to 2. The number matters: a signup that produces no outgoing mail is a broken loop that looks
like a working one.&lt;/li&gt;
&lt;li&gt;Wait for the recipient to click the link, then read the record again and watch the status move
to &lt;code&gt;Active&lt;/code&gt;, the group count move with it, and the unconfirmed count go back to zero.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Only after step 4 is the pipeline verified. Steps 1 to 3 prove the plumbing; step 4 proves the
part a reader actually performs.&lt;/p&gt;
&lt;h2 id=&quot;the-thing-that-failed-was-the-inbox&quot;&gt;The thing that failed was the inbox&lt;/h2&gt;
&lt;p&gt;The confirmation email was sent. It never arrived.&lt;/p&gt;
&lt;p&gt;The address in question was a QQ Mail account, and QQ Mail silently drops mail from foreign
senders. No bounce, no spam folder entry, nothing in the logs on my side. The mailing tool
reported a successful send, and it was telling the truth.&lt;/p&gt;
&lt;p&gt;This is worth stating plainly because it is invisible from the sending side. Every dashboard I can
look at says the email left. The only way to learn otherwise is to ask the person at the other end,
and a signup flow that depends on them telling you is a signup flow that will lose readers without
ever showing an error. The practical fixes are ordinary ones: tell people which folder to check,
use a mailbox you trust for your own tests, and expect some fraction of a list to be unreachable
through no fault of your own.&lt;/p&gt;
&lt;h2 id=&quot;two-smaller-traps&quot;&gt;Two smaller traps&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The dashboard’s copy was arguing with the site.&lt;/strong&gt; The embedded form ships with a default headline
and the line “Signup for news and special offers!”, which is roughly the opposite of what this site
promises. The form editor is a client-rendered app, and synthetic clicks from a scripted browser
do not reach its text-editing state, so there was no clean programmatic fix. I hid those two
elements in CSS and let the site’s own heading stand. That is a workaround and it is labelled as
one in the stylesheet.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The status filter is not a native control.&lt;/strong&gt; Selecting “Unconfirmed” means clicking a custom
overlay, and clicking where the option was last time lands on nothing, because the menu re-renders.
What works: send an Escape first in case the menu is already open, click the trigger, re-read the
option’s coordinates after it opens, and then click. The confirmation that the filter is applied is
a line of text reading which status is being shown, not the menu’s own state.&lt;/p&gt;
&lt;h2 id=&quot;cleaning-up-after-yourself&quot;&gt;Cleaning up after yourself&lt;/h2&gt;
&lt;p&gt;The test created subscriber records in a real account, so the last step was deleting them, which is
also the step that is easy to skip. I kept two addresses deliberately, for future sends: the one
that can receive foreign mail and the one that cannot. Knowing which is which is useful the next
time a confirmation email “sends successfully”.&lt;/p&gt;
&lt;h2 id=&quot;what-i-would-still-verify&quot;&gt;What I would still verify&lt;/h2&gt;
&lt;p&gt;Payment. The support page has a working button and a payment page that returns 200, and I have
never completed a real transaction through it. Until I do, the honest description of that button is
“untested”, not “working”. The same standard applies here: the signup was verified end to end
because a person on the other side clicked a link and the state changed. Anything less than that
is a hope with a green tick next to it.&lt;/p&gt;</content:encoded></item><item><title>What this site actually ships, and the widget I can&apos;t remove</title><link>https://notebookfield.com/posts/what-this-site-ships/</link><guid isPermaLink="true">https://notebookfield.com/posts/what-this-site-ships/</guid><description>2,130 bytes of HTML, 1 KB of my own JavaScript, no framework. Then the newsletter form loads jQuery, and the number stops being funny.</description><pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;The front page of this site is a 2,130-byte HTML document containing 91 elements. I wrote 1,022
bytes of JavaScript for it, in two inline blocks, and I have no build step that could hide
anything else. That is the whole page: text, a stylesheet, two typefaces.&lt;/p&gt;
&lt;p&gt;Then you scroll to the signup form, and jQuery 3.7.1 arrives.&lt;/p&gt;
&lt;h2 id=&quot;how-i-measured-it&quot;&gt;How I measured it&lt;/h2&gt;
&lt;p&gt;Not from the build output. Build output tells you what you shipped, not what a browser fetched.
I opened the live site in headless Chrome, attached to it over the DevTools protocol, and read
&lt;code&gt;performance.getEntriesByType(&apos;resource&apos;)&lt;/code&gt; plus the navigation timing entry.&lt;/p&gt;
&lt;p&gt;The page, from my line:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;document: 2,130 bytes encoded, 91 elements in the DOM&lt;/li&gt;
&lt;li&gt;same-origin requests: 4. Stylesheet 3,498 B, Inter 48,556 B, Newsreader 132,300 B, favicon 469 B&lt;/li&gt;
&lt;li&gt;JavaScript I wrote: 1,022 B, in two inline &lt;code&gt;&amp;#x3C;script&gt;&lt;/code&gt; blocks (255 B to set the theme before
first paint, 767 B for the toggle)&lt;/li&gt;
&lt;li&gt;there is also a 309-byte inline &lt;code&gt;&amp;#x3C;script type=&quot;application/ld+json&quot;&gt;&lt;/code&gt;, which is data for
search engines rather than code&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And the signup widget, which is a third-party embed:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;9 requests to MailerLite’s CDN&lt;/li&gt;
&lt;li&gt;225,273 bytes downloaded, of which 223,312 bytes is JavaScript: jQuery 87,532 B, the widget’s
own loader 52,683 B, an input-mask bundle 70,970 B, a form renderer 12,127 B&lt;/li&gt;
&lt;li&gt;the remaining ~2 KB is two small stylesheets&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So the form costs about 100 times my own HTML in JavaScript, on every page that shows it.&lt;/p&gt;
&lt;p&gt;Two honest caveats on those numbers. Cross-origin responses report &lt;code&gt;transferSize: 0&lt;/code&gt; unless the
server sends &lt;code&gt;Timing-Allow-Origin&lt;/code&gt;, so the widget’s bytes came from &lt;code&gt;curl&lt;/code&gt; against the six asset
URLs, not from the browser’s own accounting. And the timings came from my machine, over a tunnel:
document response started at 1,593 ms, DOM ready at 2,151 ms, and the &lt;code&gt;load&lt;/code&gt; event at 9,526 ms.
That last number is the widget too. None of these three timings are what you would see.&lt;/p&gt;
&lt;h2 id=&quot;why-the-typefaces-are-mine-to-serve&quot;&gt;Why the typefaces are mine to serve&lt;/h2&gt;
&lt;p&gt;The two font files are the heaviest thing I send you, and they are the part I am most confident
about. They are subsets: Inter at 48 KB, Newsreader at 132 KB, latin ranges only, no third-party
request, no DNS lookup to someone else’s CDN, and a 6-month &lt;code&gt;Cache-Control&lt;/code&gt; from my own host.&lt;/p&gt;
&lt;p&gt;I could have linked to a font CDN instead. I decided not to, for a reason I can state precisely:
I cannot verify that CDN from where my readers are. I also cannot verify it from where I am. When
I benchmarked a list of hosts from this machine, Google Fonts answered in about a second, which
looks like proof that nothing is blocked, except that my machine routes everything through a
tunnel: a request to a foreign host leaves from a Tokyo exit node, and only requests to Chinese
hosts leave from my actual ISP address. Two different addresses are answering me depending on
which host I ask, and both of them are reported as “direct”.&lt;/p&gt;
&lt;p&gt;That is the useful lesson here, and it cost me a wrong sentence earlier in this project. A test
that does not tell you which way the packets left is not a reachability test. So the font files
ship with the site, where the only network path involved is the one that already delivered the
page.&lt;/p&gt;
&lt;h2 id=&quot;the-failure-worth-writing-down&quot;&gt;The failure worth writing down&lt;/h2&gt;
&lt;p&gt;I styled the widget to match the site and pushed it. The button stayed black. In dark mode the
email field stayed white.&lt;/p&gt;
&lt;p&gt;MailerLite’s stylesheet writes its rules with the form’s container id and &lt;code&gt;!important&lt;/code&gt;:
&lt;code&gt;#mlb2-46888919.ml-form-embedContainer ... .ml-form-embedSubmit button { background-color: #000 !important }&lt;/code&gt;. My override was &lt;code&gt;.newsletter .ml-form-embedSubmit button&lt;/code&gt;, which has two classes.
An id plus four classes beats two classes no matter how many &lt;code&gt;!important&lt;/code&gt; flags I add, because
&lt;code&gt;!important&lt;/code&gt; only settles ties inside the same specificity tier.&lt;/p&gt;
&lt;p&gt;The fix is boring: repeat the same id in my selector and add one more class, so my rule wins on
specificity and then on &lt;code&gt;!important&lt;/code&gt;. The interesting part is that the screenshot looked fine. A
black button on a light page looks like a design choice. I only caught it because I stopped
looking at pixels and asked the browser for computed values:&lt;/p&gt;
&lt;pre class=&quot;astro-code github-dark&quot; style=&quot;background-color:#24292e;color:#e1e4e8; overflow-x: auto;&quot; tabindex=&quot;0&quot; data-language=&quot;js&quot;&gt;&lt;code&gt;&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#B392F0&quot;&gt;getComputedStyle&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;(document.&lt;/span&gt;&lt;span style=&quot;color:#B392F0&quot;&gt;querySelector&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;(&lt;/span&gt;&lt;span style=&quot;color:#9ECBFF&quot;&gt;&apos;.ml-embedded button[type=submit]&apos;&lt;/span&gt;&lt;span style=&quot;color:#E1E4E8&quot;&gt;)).backgroundColor&lt;/span&gt;&lt;/span&gt;
&lt;span class=&quot;line&quot;&gt;&lt;span style=&quot;color:#6A737D&quot;&gt;// &quot;rgb(0, 0, 0)&quot;  -&gt; my stylesheet was not applying at all&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Both rounds after that, light and dark, I asserted on computed colours, and kept the screenshots
for layout only. Pixels are for judging whether something looks right. They are a bad instrument
for finding out whether your CSS is in the cascade.&lt;/p&gt;
&lt;h2 id=&quot;what-i-would-do-about-the-220-kb&quot;&gt;What I would do about the 220 KB&lt;/h2&gt;
&lt;p&gt;Defer it. The widget only has to exist when a reader scrolls to it or clicks it, and the page can
load the embed then instead of in the &lt;code&gt;&amp;#x3C;head&gt;&lt;/code&gt;. That is the next change I want to make here, and
when it is done I will re-run the same measurement and put both numbers in this post, including
the case where it does not help.&lt;/p&gt;
&lt;p&gt;Until then the honest statement is: this site has no framework, no bundler, no client-side router,
no analytics, and about a kilobyte of JavaScript that I wrote. The one third-party script on it is
a newsletter form, and it is heavier than everything else on the page put together.&lt;/p&gt;
&lt;h2 id=&quot;reproducing-this&quot;&gt;Reproducing this&lt;/h2&gt;
&lt;p&gt;The measurement is a CDP script: open the page in headless Chrome, enable the Network domain, read
&lt;code&gt;performance.getEntriesByType(&apos;resource&apos;)&lt;/code&gt;, then &lt;code&gt;curl -o /dev/null -w &apos;%{size_download}&apos;&lt;/code&gt; the
asset URLs to get the cross-origin bytes the Resource Timing API refuses to report. If you want
to compare your own site, the part that matters is the last step. Browser-side byte counts for
third-party assets are usually 0, and it is easy to read that as “the widget is free”.&lt;/p&gt;</content:encoded></item><item><title>What goes on this site, and what does not</title><link>https://notebookfield.com/posts/what-goes-on-this-site/</link><guid isPermaLink="true">https://notebookfield.com/posts/what-goes-on-this-site/</guid><description>The scope of Field Notes, how often it updates, and why the numbers here have to come from a script.</description><pubDate>Wed, 07 Oct 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Most writing about AI tooling is written by people who watched a demo. This site is written
by someone who runs the same task book against five tools, then reads the 39 acceptance
checks line by line to find out which ones passed for the wrong reason.&lt;/p&gt;
&lt;p&gt;That difference is the whole point, and it comes with a cost: it is slower. So the rules here.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;One long post every two weeks, not three short ones a week.&lt;/strong&gt; Each post has to contain
something a reader can re-run, copy, or argue against — a task book, a scoring rule, a
failure taxonomy, the script that computed the number.&lt;/p&gt;
&lt;p&gt;The three posts published alongside this one are an opening batch. The fortnightly rhythm
starts after them.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Numbers come from scripts, not from impressions.&lt;/strong&gt; If a post says a tool passed 39 out of 39,
there is a script that produced that 39, and I will say what it could not see. Where I have not
verified something, the post says so in those words.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;No employer internals.&lt;/strong&gt; The method is mine to publish; the data is not. Anything under an
NDA, any customer ticket, and any internal number stays out.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;If it breaks on the third run, the post says so.&lt;/strong&gt; The failure mode is usually the most
useful sentence in the whole piece.&lt;/p&gt;
&lt;p&gt;If that sounds like the kind of thing you want in your inbox, the signup form is on the
&lt;a href=&quot;/&quot;&gt;front page&lt;/a&gt;. Until then, the &lt;a href=&quot;/posts/&quot;&gt;posts&lt;/a&gt; page is where everything lands.&lt;/p&gt;</content:encoded></item></channel></rss>