<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Sparse Vector]]></title><description><![CDATA[In a high dimensional space, most values are zero. I write about the ones that aren’t.]]></description><link>https://www.sparsevector.ai</link><image><url>https://substackcdn.com/image/fetch/$s_!p4kb!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff69b0776-4e25-47b1-8a95-cd66d40b7d9d_1254x1254.png</url><title>Sparse Vector</title><link>https://www.sparsevector.ai</link></image><generator>Substack</generator><lastBuildDate>Wed, 22 Jul 2026 10:51:06 GMT</lastBuildDate><atom:link href="https://www.sparsevector.ai/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Rishabh Gupta]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[sparsevector@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[sparsevector@substack.com]]></itunes:email><itunes:name><![CDATA[Rishabh Gupta]]></itunes:name></itunes:owner><itunes:author><![CDATA[Rishabh Gupta]]></itunes:author><googleplay:owner><![CDATA[sparsevector@substack.com]]></googleplay:owner><googleplay:email><![CDATA[sparsevector@substack.com]]></googleplay:email><googleplay:author><![CDATA[Rishabh Gupta]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Sensitive Data and Vector Weaknesses: What OWASP's LLM02 and LLM08 Actually Cover]]></title><description><![CDATA[Part 2 of 6 in a series going through the OWASP LLM Top 10 one category at a time &#8212; this time: what leaks, and where it lives]]></description><link>https://www.sparsevector.ai/p/the-owasp-llm-top-10-a-practitioners-af1</link><guid isPermaLink="false">https://www.sparsevector.ai/p/the-owasp-llm-top-10-a-practitioners-af1</guid><dc:creator><![CDATA[Rishabh Gupta]]></dc:creator><pubDate>Tue, 21 Jul 2026 14:40:07 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!1UIJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f247e1-ebde-48d3-99ae-240631f99bf5_1080x1400.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="https://open.substack.com/pub/sparsevector/p/the-owasp-llm-top-10-a-practitioners?r=gz99v&amp;utm_campaign=post&amp;utm_medium=web">Part 1</a> of this series covered LLM01: Prompt Injection. This post covers two categories together, because they describe the same underlying problem from two different angles: what a RAG system exposes, and where that exposure actually lives &#8212; in the vector store itself.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1UIJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f247e1-ebde-48d3-99ae-240631f99bf5_1080x1400.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1UIJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f247e1-ebde-48d3-99ae-240631f99bf5_1080x1400.png 424w, https://substackcdn.com/image/fetch/$s_!1UIJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f247e1-ebde-48d3-99ae-240631f99bf5_1080x1400.png 848w, https://substackcdn.com/image/fetch/$s_!1UIJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f247e1-ebde-48d3-99ae-240631f99bf5_1080x1400.png 1272w, https://substackcdn.com/image/fetch/$s_!1UIJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f247e1-ebde-48d3-99ae-240631f99bf5_1080x1400.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1UIJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f247e1-ebde-48d3-99ae-240631f99bf5_1080x1400.png" width="1080" height="1400" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a9f247e1-ebde-48d3-99ae-240631f99bf5_1080x1400.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1400,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:137291,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://sparsevector.substack.com/i/207887353?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f247e1-ebde-48d3-99ae-240631f99bf5_1080x1400.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1UIJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f247e1-ebde-48d3-99ae-240631f99bf5_1080x1400.png 424w, https://substackcdn.com/image/fetch/$s_!1UIJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f247e1-ebde-48d3-99ae-240631f99bf5_1080x1400.png 848w, https://substackcdn.com/image/fetch/$s_!1UIJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f247e1-ebde-48d3-99ae-240631f99bf5_1080x1400.png 1272w, https://substackcdn.com/image/fetch/$s_!1UIJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f247e1-ebde-48d3-99ae-240631f99bf5_1080x1400.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><div><hr></div><h3>LLM02: Sensitive Information Disclosure</h3><p>Per the <a href="https://genai.owasp.org/llmrisk/llm022025-sensitive-information-disclosure/">official OWASP entry</a>, sensitive information disclosure covers a broader category than &#8220;the model leaked a password.&#8221; OWASP&#8217;s own scope includes personal identifiable information, financial details, health records, confidential business data, security credentials, legal documents, and &#8212; for closed or foundation models &#8212; proprietary training methods and source code.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.sparsevector.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sparse Vector! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The document frames this as a two-directional problem. Sensitive data can leak <em>out</em> through a model&#8217;s output, or it can be absorbed <em>in</em> through training or fine-tuning on data that shouldn&#8217;t have been included in the first place. Different failure modes, different fixes.</p><p><strong>Three named patterns, one with real precedent.</strong> OWASP lists PII leakage (personal data disclosed during ordinary use, no attack required), proprietary algorithm exposure, and sensitive business data disclosure. For the second one, OWASP anchors to a concrete precedent: the <a href="https://avidml.org/database/avid-2023-v009/">&#8220;Proof Pudding&#8221; attack</a> (CVE-2019-20634), where disclosed training data let attackers reconstruct and invert a model, then use that reconstruction to bypass the email spam filter the model was meant to enforce. It&#8217;s an older, non-LLM-specific case, but it&#8217;s the hard technical precedent the category is built on.</p><p>OWASP also points to the <a href="https://cybernews.com/security/chatgpt-samsung-leak-explained-lessons/">ChatGPT/Samsung incident</a> as real-world context: employees pasted proprietary source code into ChatGPT, which then became part of what the model could potentially surface elsewhere. Nobody injected anything. Someone just typed confidential code into a chat box. That&#8217;s the clearest illustration that this category isn&#8217;t purely an attack category &#8212; it&#8217;s often just a <em>usage</em> category.</p><p><strong>Mitigations read like general data-protection hygiene:</strong> sanitize inputs before they reach training data, enforce least-privilege access to sensitive data, use privacy-preserving techniques like federated learning and differential privacy, educate users not to paste sensitive information into a chat box, and keep the system prompt concealed from override or extraction. None of this is exotic if you&#8217;ve done data protection work before LLMs existed. The novelty is that a conversational interface makes it easy to forget these protections still apply.</p><div><hr></div><h3>LLM08: Vector and Embedding Weaknesses</h3><p>This is the newest and most RAG-specific category in the entire OWASP list, added in the 2025 edition specifically because RAG had become the dominant deployment pattern. Where LLM02 is about sensitive data in general, LLM08 is about the <em>mechanism</em> &#8212; the vector store itself &#8212; that makes a RAG-specific version of that leakage possible.</p><p>Per the <a href="https://genai.owasp.org/llmrisk/llm082025-vector-and-embedding-weaknesses/">official entry</a>, the core risk is that inadequate or misaligned access controls on embeddings can let a model retrieve and disclose personal data, proprietary information, or sensitive content it technically has stored but was never supposed to surface to a given user.</p><p><strong>Three specific failure patterns are named:</strong></p><p><strong>Multi-tenant context leakage.</strong> When multiple users or applications share the same vector database, there&#8217;s a real risk that one tenant&#8217;s query surfaces another tenant&#8217;s data. This is the same access-boundary problem as any shared-infrastructure system, just applied to embeddings instead of rows in a SQL table.</p><p><strong>Embedding inversion.</strong> Attackers can exploit vulnerabilities to invert embeddings and recover meaningful amounts of the original source information &#8212; turning a vector that was supposed to be an opaque numerical representation back into readable content. This directly undermines the assumption that embeddings are a &#8220;safe,&#8221; anonymized form of the original data.</p><p><strong>Data poisoning.</strong> Poisoned data can enter a vector store intentionally (a malicious actor) or unintentionally (a bad upstream data source), manipulating what the system retrieves and, downstream, what it tells users. This overlaps with the corpus-poisoning territory from Part 1&#8217;s coverage of LLM01, but the framing here is about data integrity in storage, not injection through retrieval.</p><p><strong>How this differs from prompt injection, concretely:</strong> prompt injection targets model instructions directly. LLM08 attacks manipulate the data layer the model retrieves and implicitly trusts &#8212; often without touching a prompt at all. That distinction matters operationally: your defenses against one don&#8217;t automatically cover the other.</p><p><strong>A practical mitigation worth naming specifically:</strong> for multi-tenant leakage, one straightforward fix is separating data into per-tenant collections or indexes at the vector-store level &#8212; the same idea, at the infrastructure layer, as row-level security in a traditional database. If you&#8217;re building a RAG system that serves more than one customer or user group from a shared corpus, this is the first thing to check, not an afterthought.</p><div><hr></div><h3>What This Means If You&#8217;re Building RAG</h3><p>The pattern connecting LLM02 and LLM08 to each other, and back to Part 1&#8217;s coverage of LLM01, is the same each time: what you ingest is what you&#8217;re exposed to, and the exposure surface isn&#8217;t just the model&#8217;s output &#8212; it&#8217;s every layer of the pipeline that stores, indexes, or retrieves that data. A RAG system that ingests contracts, internal wikis, support tickets, or medical records is doing exactly what these two categories warn about, by design. The question isn&#8217;t whether that data is exposed somewhere in the system. It&#8217;s whether you&#8217;ve deliberately controlled who can retrieve it, and whether the storage layer itself &#8212; not just the chat interface &#8212; enforces that control.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.sparsevector.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sparse Vector! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[LLM01: Prompt Injection — The Three-Axis Framework Nobody Reads]]></title><description><![CDATA[Part 1 of 6 in a series going through the OWASP LLM Top 10 one category at a time]]></description><link>https://www.sparsevector.ai/p/the-owasp-llm-top-10-a-practitioners</link><guid isPermaLink="false">https://www.sparsevector.ai/p/the-owasp-llm-top-10-a-practitioners</guid><dc:creator><![CDATA[Rishabh Gupta]]></dc:creator><pubDate>Tue, 14 Jul 2026 14:55:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!1MM2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc597e183-5f4d-4488-a0b4-d4b6b5edc9ba_1080x1450.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>If you&#8217;re building anything with LLMs in production, there&#8217;s a good chance you&#8217;ve heard of the <a href="https://github.com/GenAI-Security-Project/GenAI-LLM-Top10">OWASP LLM Top 10</a> without actually reading it. It&#8217;s the closest thing the industry has to a shared vocabulary for AI security risks &#8212; but the actual document is dense, heavily cited, and easy to skim past without absorbing the parts that matter.</p><p>Here&#8217;s what&#8217;s worth knowing from the LLM01 entry specifically, especially if you&#8217;re building RAG systems, agents, or anything with retrieval.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.sparsevector.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.sparsevector.ai/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h3>Why LLM Security Is a Different Problem</h3><p>Traditional application security has a clear boundary: code is trusted, input is untrusted. SQL injection exists because you can parameterize queries to separate the two.</p><p>LLMs don&#8217;t have this boundary. The system prompt, the user&#8217;s question, retrieved documents, and tool outputs are all just tokens on the same stream. There&#8217;s no equivalent of a parameterized query &#8212; the model reads everything as potential instruction.</p><p>Three properties make this worse in real systems:</p><p><strong>Context-window pooling.</strong> Everything shares one input &#8212; system prompt, user input, RAG documents, tool outputs, memory. No enforced trust boundary between them.</p><p><strong>Memory persistence.</strong> An injection that writes to a vector store or RAG corpus doesn&#8217;t just affect one conversation &#8212; it poisons every future session that retrieves from that store.</p><p><strong>Agentic execution.</strong> When a model&#8217;s output drives tool calls &#8212; file systems, APIs, email, MCP servers &#8212; the blast radius of a successful attack extends from the chat window to everything the agent&#8217;s tools can reach.</p><div><hr></div><h3>The Three-Axis Framework</h3><p>The most practically useful thing in the entire <a href="https://github.com/GenAI-Security-Project/GenAI-LLM-Top10/blob/main/2026/final/LLM01_PromptInjection.md">OWASP LLM01:2026 document</a>, is a framework for classifying attacks along three independent axes:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1MM2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc597e183-5f4d-4488-a0b4-d4b6b5edc9ba_1080x1450.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1MM2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc597e183-5f4d-4488-a0b4-d4b6b5edc9ba_1080x1450.png 424w, https://substackcdn.com/image/fetch/$s_!1MM2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc597e183-5f4d-4488-a0b4-d4b6b5edc9ba_1080x1450.png 848w, https://substackcdn.com/image/fetch/$s_!1MM2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc597e183-5f4d-4488-a0b4-d4b6b5edc9ba_1080x1450.png 1272w, https://substackcdn.com/image/fetch/$s_!1MM2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc597e183-5f4d-4488-a0b4-d4b6b5edc9ba_1080x1450.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1MM2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc597e183-5f4d-4488-a0b4-d4b6b5edc9ba_1080x1450.png" width="1080" height="1450" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c597e183-5f4d-4488-a0b4-d4b6b5edc9ba_1080x1450.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1450,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:144810,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://sparsevector.substack.com/i/206971676?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc597e183-5f4d-4488-a0b4-d4b6b5edc9ba_1080x1450.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1MM2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc597e183-5f4d-4488-a0b4-d4b6b5edc9ba_1080x1450.png 424w, https://substackcdn.com/image/fetch/$s_!1MM2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc597e183-5f4d-4488-a0b4-d4b6b5edc9ba_1080x1450.png 848w, https://substackcdn.com/image/fetch/$s_!1MM2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc597e183-5f4d-4488-a0b4-d4b6b5edc9ba_1080x1450.png 1272w, https://substackcdn.com/image/fetch/$s_!1MM2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc597e183-5f4d-4488-a0b4-d4b6b5edc9ba_1080x1450.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p><strong>(a) Delivery surface</strong> &#8212; how the malicious content reaches the model. The document breaks this axis down into three trust levels, worth using directly to audit your own system:</p><ul><li><p><strong>Untrusted surfaces</strong> &#8212; public web pages, emails from unknown senders, public files, search results. Treat as suspicious by default.</p></li><li><p><strong>Semi-trusted surfaces</strong> &#8212; issue titles in a public bug tracker, package READMEs and changelogs, third-party API responses. You trust the platform, not necessarily every contributor on it.</p></li><li><p><strong>Trusted surfaces</strong> &#8212; code in a repository you own, rows in your own production database, internal documents, your own emails or calendar. This is the dangerous blind spot: you don&#8217;t expect an attacker here, but they can plant content via an unrelated upstream vector &#8212; a public bug-report form, a customer support ticket &#8212; that eventually lands somewhere you implicitly trust.</p></li></ul><p>The shared insight across all three: the attacker doesn&#8217;t need to compromise your backend directly. They place text somewhere your LLM will eventually read it, and your own system &#8212; operating with your own privileges &#8212; does the damage. A defense that only scrutinizes the &#8220;untrusted&#8221; tier misses attacks arriving through the trusted tier entirely.</p><p><strong>(b) Propagation</strong> &#8212; how far the attack's effect spreads: a single response, a multi-step chain across several turns, cross-session via memory or RAG, or self-replicating across multiple agents. A single injected instruction that only affects one reply is a very different risk than one that poisons a shared knowledge base for every future user.</p><p><strong>(c) Encoding</strong> &#8212; how the payload is hidden: plain text, base64, invisible Unicode characters, or steganography embedded in an image or audio file. A payload hidden inside an image's pixel data can pass through a text-only content filter untouched, since the filter never sees anything resembling suspicious text.</p><p>Every real attack is a combination of one item from each axis. A documented <a href="https://embracethered.com/blog/posts/2024/m365-copilot-prompt-injection-tool-invocation-and-data-exfil-using-ascii-smuggling/">proof-of-concept against Microsoft 365 Copilot</a> (August 2024) combined document-based delivery (a trusted-surface- adjacent channel), a single-shot tool-invocation chain, and invisible Unicode encoding &#8212; three axes, one attack.</p><p>The practical exercise worth doing: list your system&#8217;s actual input sources and sort them into the three trust levels above. Most teams have only defended the first one.</p><div><hr></div><h3>The Finding That Should Change How You Think About RAG</h3><p><strong>As few as five documents injected into a RAG corpus achieved attack success rates above 95%</strong> on standard question-answering corpora, according to <a href="https://www.usenix.org/system/files/usenixsecurity25-zou-poisonedrag.pdf">PoisonedRAG</a> (Zou et al., USENIX Security 2025). You don&#8217;t need to compromise the backend, the model weights, or the infrastructure. You just need to get five documents into whatever corpus your RAG system retrieves from.</p><p>If your RAG pipeline ingests from any source you don&#8217;t fully control &#8212; scraped web content, user uploads, shared drives, third-party APIs &#8212; this number should worry you more than almost anything else in the document.</p><div><hr></div><h3>Five Mitigations Worth Actually Implementing</h3><p>The OWASP document lists eleven mitigations. Five are worth prioritizing:</p><p><strong>Minimum permissions.</strong> Hold API credentials and sensitive operations in application code, not in the model&#8217;s context. The model should request an action; application code should decide whether to perform it.</p><p><strong>Human confirmation before irreversible actions.</strong> Any action that sends, deletes, or modifies something outside the conversation should require explicit confirmation &#8212; not because the model can&#8217;t be trusted, but because a single injected instruction shouldn&#8217;t be able to act unsupervised.</p><p><strong>The Rule of Two.</strong> If an agent has (A) untrusted input, (B) access to sensitive data, and (C) the ability to change state &#8212; require human approval before any action. Having all three simultaneously is the actual danger zone; most production agents have this without realizing it.</p><p><strong>Treat memory writes as privileged operations.</strong> Anything written to a persistent store &#8212; a vector database, a long-term memory system &#8212; should be logged with its source and, ideally, screened for embedded instructions before it&#8217;s trusted by future sessions.</p><p><strong>Pin, sign, and verify MCP servers.</strong> As agent tooling spreads through the Model Context Protocol, the tool descriptions themselves become an attack surface. A malicious tool description can hijack agent behavior before the user ever types anything.</p><div><hr></div><h3>Numbers Worth Remembering</h3><ul><li><p><strong>5</strong> documents, attack success rates <strong>above 95%</strong> &#8212; <a href="https://www.usenix.org/system/files/usenixsecurity25-zou-poisonedrag.pdf">PoisonedRAG</a>, Zou et al., USENIX Security 2025</p></li><li><p><strong>~90%</strong> attack success rate for adaptive attacks against defenses that showed near-zero success rate under static testing &#8212; <a href="https://arxiv.org/abs/2510.09023">Nasr &amp; Carlini, &#8220;The Attacker Moves Second&#8221;</a>, arXiv:2510.09023, October 2025</p></li></ul><p>A defense benchmark measuring only static attacks tells you very little about resistance to an adaptive adversary who knows the defense exists.</p><div><hr></div><h3>What This Means If You&#8217;re Building RAG</h3><p>Any RAG system that ingests from a source it doesn&#8217;t fully control &#8212; user uploads, scraped web content, shared drives, third-party APIs &#8212; is vulnerable to what the document calls indirect prompt injection and RAG repository poisoning, both covered directly in this entry.</p><p>The uncomfortable pattern across the document&#8217;s mitigations is how little of the burden falls on the model itself. Almost every effective defense lives in the application layer &#8212; permission boundaries, human confirmation, provenance tracking on what gets ingested. The model is not going to solve this problem for you.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.sparsevector.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Sparse Vector! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[I Learned Rust by Rebuilding My RAG Pipeline — Here's What Python Was Hiding From You]]></title><description><![CDATA[The standard Rust advice: read the book, do the exercises, build a todo app.]]></description><link>https://www.sparsevector.ai/p/i-learned-rust-by-rebuilding-my-rag-pipeline-heres-what-python-was-hiding-from-you</link><guid isPermaLink="false">https://www.sparsevector.ai/p/i-learned-rust-by-rebuilding-my-rag-pipeline-heres-what-python-was-hiding-from-you</guid><dc:creator><![CDATA[Rishabh Gupta]]></dc:creator><pubDate>Tue, 07 Jul 2026 15:53:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!p4kb!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff69b0776-4e25-47b1-8a95-cd66d40b7d9d_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The standard Rust advice: read the book, do the exercises, build a todo app. I tried something different &#8212; I took a RAG pipeline I'd already built in Python and rebuilt the core of it in Rust. Same operations, same payload schema, same re-ranking logic. Familiar problem, unfamiliar language.</p><p>The bet: learning is faster when you already know what you're trying to build.</p><p>The pipeline was the retrieval layer from <a href="https://asks1.com/">AskS1.com</a> &#8212; a tool for querying the SpaceX S-1 filing. It embeds a user's question, searches a Qdrant vector store for the most relevant chunks, applies a re-ranking penalty to summary pages, and returns the top results. About 50 lines of Python. I rebuilt it in Rust across 8 hours.</p><p>Here's what the borrow checker taught me that 4 years of Python didn't.</p><div><hr></div><h3>The Three Things Python Was Hiding</h3><h4>1. String ownership</h4><p>In Python, passing a string to a function and using it afterward is unremarkable:</p><p>python</p><pre><code>def takes(s):
    print(s)

s = "hello"
takes(s)
print(s)  # works fine</code></pre><p>The same code in Rust is a compile error:</p><p>rust</p><pre><code>fn takes(s: String) {
    println!("{}", s);
}

let s = String::from("hello");
takes(s);
println!("{}", s);  // error: value borrowed after move</code></pre><p>The error message:</p><pre><code>value borrowed here after move
move occurs because `s` has type `String`, 
which does not implement the `Copy` trait</code></pre><p>What Python was hiding: when you pass a String to a function, ownership transfers. The original variable is gone. Python's garbage collector tracks references automatically &#8212; you never see this. Rust makes it explicit at compile time.</p><p>The fix is either to clone (takes(s.clone())) or borrow (takes(&amp;s)). Clone creates a new heap allocation. Borrow passes a reference &#8212; no allocation, no ownership transfer, cheaper. In Python, every string argument is effectively a borrow without you knowing it. In Rust, you choose.</p><p>This took 20 minutes to understand. It will probably change how I read Python code for the rest of my career.</p><div><hr></div><h4>2. Mutation while borrowed</h4><p>Python lets you do this without complaint:</p><p>python</p><pre><code>v = [1, 2, 3]
first = v[0]   # get a reference to the first element
v.append(4)    # mutate the list
print(first)   # still works</code></pre><p>Rust blocks it:</p><p>rust</p><pre><code>let mut v = vec![1, 2, 3];
let first = &amp;v[0];
v.push(4);        // error: cannot borrow `v` as mutable
                  // because it is also borrowed as immutable
println!("{}", first);</code></pre><p>At first this feels like the borrow checker being pedantic. It isn't. The reason: push might reallocate the Vec's memory if it grows beyond its current capacity. If that happens, first would be pointing at freed memory &#8212; a use-after-free bug. Python's runtime handles this by tracking all references and preventing deallocation. Rust catches it at compile time with zero runtime cost.</p><p>The fix: use first before mutating.</p><p>rust</p><pre><code>let mut v = vec![1, 2, 3];
let first = &amp;v[0];
println!("{}", first);  // use borrow here &#8212; ends after this line
v.push(4);              // now safe</code></pre><p>What Python was hiding: every time you modify a list while holding a reference to one of its elements, Python's runtime is doing work to keep that safe. Rust eliminates that work by making the constraint explicit.</p><div><hr></div><h4>3. Reference vs. value</h4><p>This one is subtle and shows up constantly in real Qdrant code.</p><p>When you iterate over a Vec of tuples, the iterator yields references to each tuple. If the tuple contains an integer, destructuring gives you a reference to that field &#8212; &amp;i64 &#8212; not the integer itself:</p><p>rust</p><pre><code>let chunks = vec![
    ("Starlink revenue was $11.4B", 89_i64),
    ("SpaceX launched 165 times in 2025", 125_i64),
];

for (text, page) in chunks.iter() {
    // page is &amp;i64 here, not i64 &#8212; iterating with .iter() borrows
    // each tuple, and destructuring binds a reference to each field
    // rather than moving the value out. To use page as a plain i64,
    // dereference it with *page.
    println!("Page: {}", *page);
}</code></pre><p>In Python this is invisible &#8212; iteration just gives you the value. In Rust, .iter() borrows rather than moves, so destructuring a tuple reference still yields references to its fields, not owned values. Forgetting the <code>*</code> when you need the actual value is a type error:</p><pre><code>error: expected `i64`, found `&amp;i64`</code></pre><p>This appeared in the actual Qdrant payload code when building the re-ranking function. Payload values come back from the vector store as references, and extracting integers requires pattern matching through the reference layer:</p><p>rust</p><pre><code>let page = point.payload
    .get("page")
    .and_then(|v| match &amp;v.kind {
        Some(Kind::IntegerValue(i)) =&gt; Some(*i),  // *i dereferences &amp;i64 to i64
        _ =&gt; None,
    })
    .unwrap_or(0);</code></pre><p>Python never shows you this. The interpreter handles dereferencing automatically at every step. Rust shows you every level of indirection explicitly, which is verbose but makes memory layout legible.</p><div><hr></div><h3>The Same RAG Operation in Both Languages</h3><p>Here's the core of the <code>retrieve()</code> function in Python &#8212; the version running on AskS1:</p><p>python</p><pre><code>def retrieve(query, top_k=5):
    query_vector = embed_model.encode([query])[0].tolist()
    
    results = qdrant.query_points(
        collection_name=COLLECTION,
        query=query_vector,
        limit=15
    )
    
    def score(r):
        page = r.payload.get('page', 0)
        penalty = 0.15 if page &lt; 30 else 0
        return r.score - penalty
    
    reranked = sorted(results.points, key=score, reverse=True)
    return [
        {"text": r.payload["text"], "page": r.payload["page"]}
        for r in reranked[:top_k]
    ]</code></pre><p>The Rust equivalent:</p><p>rust</p><pre><code>async fn retrieve(
    client: &amp;Qdrant,
    query_vector: Vec&lt;f32&gt;,
) -&gt; Result&lt;Vec&lt;(f32, i64, String)&gt;, Box&lt;dyn std::error::Error&gt;&gt; {
    
    let results = client.query(
        QueryPointsBuilder::new("rust_warmup")
            .query(query_vector)
            .limit(15)
            .with_payload(true)
    ).await?;

    let mut scored: Vec&lt;(f32, i64, String)&gt; = results.result
        .iter()
        .map(|point| {
            let page = point.payload
                .get("page")
                .and_then(|v| match &amp;v.kind {
                    Some(Kind::IntegerValue(i)) =&gt; Some(*i),
                    _ =&gt; None,
                })
                .unwrap_or(0);

            let text = point.payload
                .get("text")
                .and_then(|v| match &amp;v.kind {
                    Some(Kind::StringValue(s)) =&gt; Some(s.clone()),
                    _ =&gt; None,
                })
                .unwrap_or_default();

            let penalty = if page &lt; 30 { 0.15 } else { 0.0 };
            (point.score - penalty, page, text)
        })
        .collect();

    scored.sort_by(|a, b| b.0.partial_cmp(&amp;a.0).unwrap());
    Ok(scored.into_iter().take(5).collect())
}</code></pre><p>The Rust version is twice as long. Every line is explicit about what it does &#8212; ownership of query_vector transfers into the builder, payload values are pattern-matched through their type variants, the sort is explicit about comparison direction.</p><p>The Python version is more readable. The Rust version has no hidden allocations, no silent failures, and the compiler guarantees the payload extraction logic is correct before the program ever runs.</p><p>Neither is better. They're optimized for different things.</p><div><hr></div><h3>Where Both Languages Converge</h3><p>The actual business logic is identical:</p><p>python</p><pre><code># Python
penalty = 0.15 if page &lt; 30 else 0.0
score = similarity - penalty</code></pre><p>rust</p><pre><code>// Rust
let penalty = if page &lt; 30 { 0.15 } else { 0.0 };
let adjusted = score - penalty;</code></pre><p>Once you get past the ownership layer, the domain logic looks the same. Rust's complexity is front-loaded &#8212; it makes you think about memory once, at the type system level, so you never think about it again at runtime. Python defers that complexity to the garbage collector, which handles it silently every time your code runs.</p><p>Neither approach is free. Python pays at runtime. Rust pays at development time.</p><div><hr></div><h3>Should Python AI Engineers Learn Rust?</h3><p>Honest answer: not for application code. Python wins on velocity, ecosystem, and AI tooling integration for the foreseeable future. If you're building a RAG pipeline, a fine-tuning script, or an agent &#8212; use Python.</p><p>But two cases where it's worth the investment:</p><p><strong>You're building infrastructure.</strong> Vector databases, embedding pipelines, inference servers. Qdrant is Rust. The new generation of AI agent frameworks &#8212; Rig, AutoAgents, OpenFANG &#8212; are Rust. If you're building the engine others build on, Rust's performance and memory safety matter in ways Python can't match. Per the <a href="https://blog.rust-lang.org/2026/03/02/2025-State-Of-Rust-Survey-results/?ref=rishabh.fyi">2025 State of Rust Survey</a> (as reported by <a href="https://thenewstack.io/rust-enterprise-developers/?ref=rishabh.fyi">The New Stack</a>), 48.8% of organizations now report non-trivial Rust usage in production, up from 38.7% in 2023.The language is no longer experimental.</p><p><strong>You want to understand what Python is doing.</strong> The borrow checker makes explicit what Python's GC hides. Even if you never ship Rust in production, 8 hours with the borrow checker will change how you think about Python memory, reference semantics, and mutation. The three things Python was hiding &#8212; ownership transfer, aliasing rules, reference indirection &#8212; are things Python engineers benefit from understanding, even if they never write a line of Rust again.</p><p>The "should I learn Rust?" question is slowly becoming "when should I learn Rust?" The answer for most Python AI engineers is probably: not now, but sooner than you think.</p>]]></content:encoded></item><item><title><![CDATA[The Algorithm Behind Every Vector Database Search — And Why It Matters for AI Engineers]]></title><description><![CDATA[When I built AskS1.com &#8212; a tool that lets you ask questions about the SpaceX S-1 filing and get cited answers &#8212; I spent a lot of time thinking about retrieval.]]></description><link>https://www.sparsevector.ai/p/the-algorithm-behind-every-vector-database-search-and-why-it-matters-for-ai-engineers</link><guid isPermaLink="false">https://www.sparsevector.ai/p/the-algorithm-behind-every-vector-database-search-and-why-it-matters-for-ai-engineers</guid><dc:creator><![CDATA[Rishabh Gupta]]></dc:creator><pubDate>Sun, 28 Jun 2026 23:58:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a40f3c44-0d95-493e-a104-f572f7173ef9_434x446.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When I built <a href="https://asks1.com/">AskS1.com</a> &#8212; a tool that lets you ask questions about the SpaceX S-1 filing and get cited answers &#8212; I spent a lot of time thinking about retrieval. How do you find the right chunks of text from a 395-page document, fast enough that someone will actually wait for the answer?</p><p>The answer turned out to be an algorithm I'd been using without fully understanding: <strong>HNSW &#8212; Hierarchical Navigable Small World graphs</strong>. It's the engine inside <a href="https://qdrant.tech/?ref=rishabh.fyi">Qdrant</a>, <a href="https://weaviate.io/?ref=rishabh.fyi">Weaviate</a>, <a href="https://pinecone.io/?ref=rishabh.fyi">Pinecone</a>, <a href="https://milvus.io/?ref=rishabh.fyi">Milvus</a> and most modern vector databases. If you're building anything with RAG, embeddings, or semantic search, HNSW is quietly doing the hardest part for you.</p><p>This post is what I wish I'd read before building AskS1.</p><div><hr></div><h3>What's a Vector Database, and Why Do You Need One?</h3><p>Before HNSW makes sense, you need the problem it solves.</p><p>Modern AI applications &#8212; RAG systems, semantic search, recommendation engines &#8212; work by converting text (or images, or audio) into vectors: arrays of floating-point numbers that represent meaning. Two pieces of text that mean similar things will have vectors that are numerically close to each other. "Starlink revenue in 2025" and "Connectivity segment financial results" will be neighbors in vector space even though they share no words.</p><p>A <strong>vector database</strong> stores these vectors and answers one question efficiently: <em>given a query vector, which stored vectors are most similar?</em> That's the retrieval step in RAG &#8212; embed the user's question, find the most similar chunks, feed them to the language model.</p><p>The naive approach is obvious: compare the query vector against every stored vector, rank by similarity, return the top K. This works fine at 1,000 vectors. At 1,000,000 vectors, it's too slow. At 100,000,000 vectors (the scale of production recommendation systems), it's completely infeasible.</p><p>This is the problem HNSW was designed to solve.</p><div><hr></div><h3>The Algorithm: How HNSW Actually Works</h3><p>The <a href="https://arxiv.org/pdf/1603.09320?ref=rishabh.fyi">original HNSW paper</a> was published in 2016 by Malkov and Yashunin. The core idea is elegant enough to explain in three paragraphs.</p><p>Three-layer HNSW graph showing nodes connected at different scales, with a search path descending from Layer 2 to Layer 0 to find the nearest neighbor to a query vector. Layer 2 sparse &#183; long-range &#183; coarse search Layer 1 medium range &#183; refined candidates Layer 0 dense &#183; short-range &#183; full precision A B C A D B E C A F D G B E H C &#9733; Q Search path Node in layer Nearest neighbor &#9733; Query Q</p><p>The amber path shows HNSW search: enter at Layer 2 (A&#8594;B), drop to Layer 1 (B&#8594;E), drop to Layer 0, find the nearest neighbor (&#9733;) to the query vector (Q).</p><p><strong>The structure.</strong> HNSW builds a multi-layer graph over your stored vectors. Each vector is a node. Nodes are connected to their nearest neighbors, but the connections are separated by scale across layers:</p><pre><code>Layer 2 (top)  &#8212; sparse, long-range connections &#8212; coarse "zoom out"
Layer 1        &#8212; medium-range connections
Layer 0 (base) &#8212; dense, short-range connections &#8212; full precision</code></pre><p>Only a small fraction of vectors appear at the top layers. Every vector appears at Layer 0.</p><p><strong>The search.</strong> When a query arrives, search starts at the top layer. The algorithm greedily hops toward the query &#8212; at each step, it moves to whichever neighbor is closest to the query vector. When it can't get any closer (local minimum), it drops to the next layer and repeats, starting from where it stopped. By the time it reaches Layer 0, it's already in the right neighborhood and finds the true nearest neighbors quickly.</p><p><strong>Why this is fast.</strong> Without the hierarchy, you'd need to scan many nodes to find the right neighborhood. The hierarchy acts like a map zoom: start at country level to find the right region, zoom to city level to find the right neighborhood, then zoom to street level to find the exact address. Each zoom-in starts from a much better position than random.</p><p>The result: <strong>logarithmic complexity</strong> &#8212; O(log n) search instead of O(n). At 1 million vectors, that's roughly 20 hops instead of 1,000,000 comparisons.</p><div><hr></div><h3>The Name Unpacked</h3><p>"Hierarchical Navigable Small World" is a mouthful. Each word earns its place:</p><p><strong>Hierarchical</strong> &#8212; the multi-layer structure that gives it logarithmic complexity. Without this, you get NSW (the predecessor algorithm), which is only polylogarithmic &#8212; still too slow at scale.</p><p><strong>Navigable</strong> &#8212; greedy routing through the graph converges to the right answer. Not all graphs have this property. The specific way HNSW constructs edges ensures that following the "closest neighbor at each step" rule actually leads you somewhere useful.</p><p><strong>Small World</strong> &#8212; any two nodes in the graph can be reached from each other in a small number of hops, regardless of graph size. This is the same "six degrees of separation" phenomenon studied in social network theory &#8212; Milgram's famous experiment. HNSW deliberately engineers this property into its graph structure.</p><div><hr></div><h3>Why Previous Approaches Failed</h3><p>It helps to understand what HNSW replaced:</p><p><strong>Brute-force / flat search</strong> &#8212; compare the query against every vector. O(n) &#8212; too slow at scale. Still used for tiny collections where speed doesn't matter.</p><p><strong>kd-trees</strong> &#8212; the classic algorithm for nearest neighbor search in low-dimensional spaces. Works well up to maybe 20 dimensions. Above that, the "curse of dimensionality" kicks in: the tree structure degrades and you end up scanning most of the tree anyway. Modern embedding models produce 384 to 1536-dimensional vectors &#8212; kd-trees are useless here.</p><p><strong>Locality-sensitive hashing (LSH)</strong> &#8212; hash similar vectors to the same bucket, search within the bucket. Works, but requires tuning many parameters and tends to need high memory for good recall.</p><p><strong>NSW (non-hierarchical)</strong> &#8212; the direct predecessor to HNSW. Good idea, but polylogarithmic complexity: as the dataset grows, each search requires evaluating an increasingly large number of nodes. HNSW's hierarchy adds a second log factor that brings it to true logarithmic scaling.</p><div><hr></div><h3>The Parameters You'll Actually Configure</h3><p>When you use a vector database like Qdrant, you don't implement HNSW &#8212; you configure it. Three parameters matter:</p><p><strong>M</strong> &#8212; the number of connections per node per layer. Higher M means better recall (more neighbors to navigate through) at the cost of more memory and slower index build time. Qdrant's default is 16, which works well for most embedding dimensions.</p><p><strong>efConstruction</strong> &#8212; the candidate list size during index building. Higher values build a better quality index at the cost of build time. Default is 100.</p><p><strong>ef</strong> &#8212; the candidate list size during search. This is the one you actually tune at query time. Higher ef means HNSW explores more candidates before returning results &#8212; better recall, slightly slower. If your RAG system is missing relevant chunks, increasing ef (or the equivalent limit parameter in your vector database client) is the first thing to try.</p><p>In AskS1, the <code>retrieve()</code> function passes <code>limit=15</code> to Qdrant's query_points(). HNSW finds 15 candidates, then a re-ranking step applies a penalty to summary pages and returns the top 5. Increasing the limit to 20-25 would give HNSW more candidates to work with &#8212; potentially improving the quality of retrieved chunks at marginal latency cost for an 871-chunk collection.</p><div><hr></div><h3>One More Idea Worth Understanding: The Neighbor Selection Heuristic</h3><p>The paper introduces a heuristic for selecting which nodes to connect during index construction that's worth understanding if you're working with clustered data.</p><p>The naive approach connects each new node to its M closest existing neighbors. This works most of the time, but fails on highly clustered data: if all your nearest neighbors are in the same cluster, you have no long-range connections to other clusters. The retriever gets stuck.</p><p>HNSW's heuristic (Figure 2 in the paper, also shown below) deliberately selects diverse neighbors &#8212; it prefers candidates that extend connectivity in new directions, even if they're not the absolute closest. The result is a graph that maintains global connectivity across clusters.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pMu6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a2828e9-2e97-4db7-827e-d90ad31cac92_434x446.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pMu6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a2828e9-2e97-4db7-827e-d90ad31cac92_434x446.png 424w, https://substackcdn.com/image/fetch/$s_!pMu6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a2828e9-2e97-4db7-827e-d90ad31cac92_434x446.png 848w, https://substackcdn.com/image/fetch/$s_!pMu6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a2828e9-2e97-4db7-827e-d90ad31cac92_434x446.png 1272w, https://substackcdn.com/image/fetch/$s_!pMu6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a2828e9-2e97-4db7-827e-d90ad31cac92_434x446.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pMu6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a2828e9-2e97-4db7-827e-d90ad31cac92_434x446.png" width="434" height="446" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6a2828e9-2e97-4db7-827e-d90ad31cac92_434x446.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:446,&quot;width&quot;:434,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!pMu6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a2828e9-2e97-4db7-827e-d90ad31cac92_434x446.png 424w, https://substackcdn.com/image/fetch/$s_!pMu6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a2828e9-2e97-4db7-827e-d90ad31cac92_434x446.png 848w, https://substackcdn.com/image/fetch/$s_!pMu6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a2828e9-2e97-4db7-827e-d90ad31cac92_434x446.png 1272w, https://substackcdn.com/image/fetch/$s_!pMu6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a2828e9-2e97-4db7-827e-d90ad31cac92_434x446.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This matters for RAG specifically. In a filing like the SpaceX S-1, chunks about "Starlink revenue" form a dense cluster. So do chunks about "governance" and "risk factors." Without cross-cluster connectivity, a query about "Elon Musk's voting power and its revenue implications" might only retrieve governance chunks, missing the revenue context entirely. HNSW's heuristic makes cross-cluster retrieval work.</p><div><hr></div><h3>Reading the Paper</h3><p>The <a href="https://arxiv.org/pdf/1603.09320?ref=rishabh.fyi">HNSW paper</a> is accessible without a deep algorithms background if you read it selectively. The sections worth your time:</p><ul><li><p><strong>Abstract</strong> &#8212; the entire algorithm in 15 lines</p></li><li><p><strong>Section 1</strong> &#8212; why naive search fails; no math required</p></li><li><p><strong>Section 3</strong> &#8212; the zoom-out/zoom-in intuition; the best explanatory section</p></li><li><p><strong>Figure 1</strong> &#8212; the layered structure, visually</p></li><li><p><strong>Figure 2</strong> &#8212; the neighbor selection heuristic</p></li><li><p><strong>Algorithm 5</strong> &#8212; the actual search procedure, only 8 lines</p></li><li><p><strong>Section 4.1</strong> &#8212; what M, mL, and efConstruction actually control</p></li></ul><p>Skip Section 2 (prior work survey), Algorithms 1-4 (implementation detail), and the experiments section (the finding is just "HNSW wins"). The math-heavy parts aren't necessary for understanding how to use it.</p><p>Total reading time at this depth: 45-60 minutes.</p><div><hr></div><h3>Why This Matters for AI Engineers</h3><p>Vector databases are now a standard component in AI engineering &#8212; RAG pipelines, semantic search, recommendation systems, and anything using embeddings routes through one. Understanding HNSW doesn't mean you'll implement it (you won't &#8212; Qdrant, Weaviate, and others handle that). But it tells you:</p><ul><li><p>Why <code>limit</code> in your vector database query is actually a recall parameter, not just a count</p></li><li><p>Why retrieval quality degrades on clustered data and what to do about it</p></li><li><p>Why building a 10-million-vector index takes a long time but search is fast</p></li><li><p>What tradeoffs you're making when you adjust M and efConstruction</p></li></ul><p>The algorithm was published in 2016. It's been powering production systems for almost a decade. If you're building with vector databases in 2026, it's worth spending an hour understanding what's actually happening when you call query_points().</p>]]></content:encoded></item><item><title><![CDATA[Claude Haiku vs Local Models: The Real Tradeoff]]></title><description><![CDATA[27.6 seconds vs 2.8 seconds.]]></description><link>https://www.sparsevector.ai/p/claude-haiku-vs-local-models-the-real-tradeoff</link><guid isPermaLink="false">https://www.sparsevector.ai/p/claude-haiku-vs-local-models-the-real-tradeoff</guid><dc:creator><![CDATA[Rishabh Gupta]]></dc:creator><pubDate>Tue, 16 Jun 2026 07:33:46 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1ad599d0-fa73-4dd8-861c-4f2b1e1cf401_1760x858.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BNF7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c11f78b-9b42-4b73-8077-1bb5efdb8f9b_1760x858.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BNF7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c11f78b-9b42-4b73-8077-1bb5efdb8f9b_1760x858.png 424w, https://substackcdn.com/image/fetch/$s_!BNF7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c11f78b-9b42-4b73-8077-1bb5efdb8f9b_1760x858.png 848w, https://substackcdn.com/image/fetch/$s_!BNF7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c11f78b-9b42-4b73-8077-1bb5efdb8f9b_1760x858.png 1272w, https://substackcdn.com/image/fetch/$s_!BNF7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c11f78b-9b42-4b73-8077-1bb5efdb8f9b_1760x858.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BNF7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c11f78b-9b42-4b73-8077-1bb5efdb8f9b_1760x858.png" width="1760" height="858" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9c11f78b-9b42-4b73-8077-1bb5efdb8f9b_1760x858.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:858,&quot;width&quot;:1760,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!BNF7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c11f78b-9b42-4b73-8077-1bb5efdb8f9b_1760x858.png 424w, https://substackcdn.com/image/fetch/$s_!BNF7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c11f78b-9b42-4b73-8077-1bb5efdb8f9b_1760x858.png 848w, https://substackcdn.com/image/fetch/$s_!BNF7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c11f78b-9b42-4b73-8077-1bb5efdb8f9b_1760x858.png 1272w, https://substackcdn.com/image/fetch/$s_!BNF7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c11f78b-9b42-4b73-8077-1bb5efdb8f9b_1760x858.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>27.6 seconds vs 2.8 seconds. That gap isn't a benchmark footnote &#8212; it's the difference between a product people use and one they abandon.</p><p>I was <a href="https://sparsevector.substack.com/p/how-i-built-a-rag-system-on-the-spacex-s-1-in-one-weekend">building</a> <a href="https://asks1.com/">AskS1.com</a>, a RAG system for querying the SpaceX S-1. The generation step &#8212; taking retrieved chunks and producing a cited answer &#8212; needed to be fast enough that someone would actually wait for it. I benchmarked five models to find out which one earned that spot: Claude Haiku, and four 7-14B local models running on a Mac Mini M4 via Ollama.</p><p>The overall numbers looked like a rounding error. The category breakdown told a different story.</p><div><hr></div><h3>How I Evaluated</h3><p>15 questions across three categories, same retrieved context for every model.</p><p><strong>Factual recall</strong> &#8212; can the model extract a specific number correctly?</p><pre><code>"What is SpaceX's total revenue for 2025?"
"How many Starlink subscribers does SpaceX have as of Q1 2026?"
"What is SpaceX's total debt as of Q1 2026?"</code></pre><p><strong>Multi-step reasoning</strong> &#8212; does the model connect information across sections and form a judgment?</p><pre><code>"Why is SpaceX's AI segment consuming 62-76% of capex but generating 
only 17% of revenue? Is this a concern?"
"Why can't Elon Musk be removed as CEO without his own approval?"
"How does SpaceX's vertical integration give it an advantage?"</code></pre><p><strong>Structured output</strong> &#8212; can the model follow formatting instructions precisely?</p><pre><code>"Summarize SpaceX's three business segments in a markdown table 
with columns: Segment, Revenue, Operating Income, Key Product."
"List the top 5 risk factors in order of severity."
"Summarize Elon Musk's compensation structure in exactly 4 bullets."</code></pre><p>One factual question was a deliberate curveball &#8212; "What RL algorithm does DeepSeek use?" &#8212; unrelated to SpaceX entirely, testing whether models would admit "I don't know" or hallucinate an answer just because the context was about a tech company.</p><div><hr></div><h3>Scoring</h3><p>Two methods for two question types.</p><p><strong>Factual recall</strong> &#8212; scored against ground-truth figures pulled directly from the filing. Exact numbers, keyword matching &#8212; does the answer contain the correct revenue figure, subscriber count, debt number.</p><p><strong>Reasoning and structured output</strong> &#8212; scored 1-5 by Claude Sonnet as an LLM judge, evaluating coherence, accuracy, and instruction-following. These don't have single correct answers &#8212; "is this sustainable?" requires judgment, not pattern matching.</p><div><hr></div><h3>The Results</h3><pre><code>Model            Overall  Factual  Reasoning  Structured  Latency
claude-haiku        4.7     5.0       4.8        4.4       2.8s
phi4:14b            4.5     4.4       4.5        4.6      27.6s
qwen2.5:14b         4.4     4.4       4.2        4.6      26.9s
mistral:7b          4.4     4.4       4.0        4.6       9.0s
deepseek-r1:14b     4.3     4.4       3.8        4.6     102.8s</code></pre><p>A 0.2-4.4 point spread on a 5-point scale looks like noise. It isn't &#8212; it's three different stories stacked on top of each other.</p><div><hr></div><h3>Where the Gap Actually Lives</h3><p><strong>Structured output: local models win.</strong> Every local model scored 4.6, ahead of Haiku's 4.4. Following "exactly 4 bullets" or "markdown table with these columns" doesn't require deep reasoning, and the local models were if anything slightly more literal about compliance.</p><p><strong>Reasoning: this is where the real gap is.</strong> Haiku scored 4.8. deepseek-r1:14b scored 3.8 &#8212; a full point lower, despite taking 102.8 seconds per question, 37x Haiku's latency. These questions asked models to connect numbers across sections and form a judgment &#8212; "ARPU is declining but revenue is growing &#8212; is this sustainable, and why?" This is where size and training quality actually show up. Interestingly, phi4:14b (4.5) and qwen2.5:14b (4.2) &#8212; both 14B &#8212; outperformed deepseek-r1:14b (3.8) despite being the same size class. Reasoning quality isn't just a parameter-count story.</p><p><strong>Factual recall: one question did almost all the damage.</strong> Four of five factual questions, every model scored a perfect 5.0. The entire gap traces to one question &#8212; <em>"How many Starlink subscribers does SpaceX have as of Q1 2026?"</em> All four local models answered "10,300 thousand (or 10.3 million)" &#8212; numerically correct, but the "10,300 thousand" phrasing tripped the keyword scorer. Haiku said "10.3 million" cleanly and scored full marks. Not a knowledge gap. A units-formatting quirk that cost 2.3 points on one question out of fifteen.</p><p>So the honest summary: for structured tasks, local models are competitive or better. For reasoning, there's a real gap, and it scales with model quality more than raw size. For factual recall, the "gap" was mostly an artifact of how I scored one question.</p><p>(And for the DeepSeek curveball &#8212; Haiku, phi4, qwen2.5, and deepseek-r1 all correctly said "I don't know." mistral:7b confidently described "DeepSeak, a spacecraft navigation autonomous docking system developed by SpaceX" &#8212; a system that does not exist. A small reminder that "I don't know" is sometimes the only correct answer, and not every model knows that.)</p><div><hr></div><h3>The Cost Angle</h3><p>Estimating cost per query for both:</p><p><strong>Claude Haiku</strong> &#8212; roughly 2,900 input tokens (context + system prompt + question) and ~400 output tokens per query comes to about <strong>$0.004 per query</strong>.</p><p><strong>Mac Mini M4 electricity</strong> &#8212; 27.6 seconds at ~25W draw works out to about <strong>$0.00006 per query</strong> &#8212; roughly 65x cheaper than the API call, in pure electricity terms.</p><p>Neither number matters at the scale of a side project. The Mac Mini is "free" because I already own it. The API cost is "free" because it's a fraction of a cent. Cost only becomes the deciding factor at high query volume &#8212; thousands of requests per day, where $0.004 &#215; 10,000 = $40/day starts to add up against hardware you already paid for once.</p><div><hr></div><h3>So When Do Local Models Make Sense?</h3><p>Not "Haiku wins, always." Local models make sense when:</p><ul><li><p><strong>Privacy matters</strong> &#8212; documents that can't leave your machine</p></li><li><p><strong>Offline access is required</strong> &#8212; no network dependency</p></li><li><p><strong>Volume is high enough</strong> that per-query API cost compounds meaningfully</p></li><li><p><strong>Latency tolerance is high</strong> &#8212; batch processing, overnight jobs, anything where 27 seconds vs 2.8 seconds doesn't matter to a human waiting</p></li></ul><p>For <a href="https://asks1.com/">AskS1</a> &#8212; a public tool where someone types a question and waits &#8212; 2.8 seconds is the only viable answer. But for the private Google Drive knowledge base I built on the same Mac Mini, the calculus flips entirely: nothing leaves my machine, nobody's waiting in real time, and the documents are mine. Local models there aren't a compromise &#8212; they're the right tool.</p>]]></content:encoded></item><item><title><![CDATA[How I Built a RAG System on the SpaceX S-1 in One Weekend]]></title><description><![CDATA[SpaceX filed a 389-page S-1 on May 20, 2026.]]></description><link>https://www.sparsevector.ai/p/how-i-built-a-rag-system-on-the-spacex-s-1-in-one-weekend</link><guid isPermaLink="false">https://www.sparsevector.ai/p/how-i-built-a-rag-system-on-the-spacex-s-1-in-one-weekend</guid><dc:creator><![CDATA[Rishabh Gupta]]></dc:creator><pubDate>Wed, 10 Jun 2026 15:54:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ac64c39e-1294-4e0e-a5ed-f4fe52a49749_1887x1001.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>SpaceX filed a 389-page S-1 on May 20, 2026. I read the news, opened the SEC EDGAR filing, and immediately hit the same wall everyone hits &#8212; 389 pages of dense legal and financial disclosure, no search, no way to ask a direct question and get a cited answer.</p><p>The summaries floating around were useful for headlines. Useless for anything specific. "SpaceX is profitable" tells you nothing about which segments are driving it, what the margin trajectory looks like, or what governance risks the company is flagging. For that, you need the actual text, with a page reference you can verify.</p><p>So I built <a href="https://asks1.com/">AskS1.com</a>. Here's what that actually involved &#8212; including the parts that didn't work.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!iiV7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59186c89-d148-4904-86fd-41cb23bf0307_1887x1001.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!iiV7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59186c89-d148-4904-86fd-41cb23bf0307_1887x1001.png 424w, https://substackcdn.com/image/fetch/$s_!iiV7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59186c89-d148-4904-86fd-41cb23bf0307_1887x1001.png 848w, https://substackcdn.com/image/fetch/$s_!iiV7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59186c89-d148-4904-86fd-41cb23bf0307_1887x1001.png 1272w, https://substackcdn.com/image/fetch/$s_!iiV7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59186c89-d148-4904-86fd-41cb23bf0307_1887x1001.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!iiV7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59186c89-d148-4904-86fd-41cb23bf0307_1887x1001.png" width="1887" height="1001" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/59186c89-d148-4904-86fd-41cb23bf0307_1887x1001.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1001,&quot;width&quot;:1887,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!iiV7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59186c89-d148-4904-86fd-41cb23bf0307_1887x1001.png 424w, https://substackcdn.com/image/fetch/$s_!iiV7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59186c89-d148-4904-86fd-41cb23bf0307_1887x1001.png 848w, https://substackcdn.com/image/fetch/$s_!iiV7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59186c89-d148-4904-86fd-41cb23bf0307_1887x1001.png 1272w, https://substackcdn.com/image/fetch/$s_!iiV7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59186c89-d148-4904-86fd-41cb23bf0307_1887x1001.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KTzi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82d62528-81aa-4c34-8a48-e324ac35a66c_1887x1000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KTzi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82d62528-81aa-4c34-8a48-e324ac35a66c_1887x1000.png 424w, https://substackcdn.com/image/fetch/$s_!KTzi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82d62528-81aa-4c34-8a48-e324ac35a66c_1887x1000.png 848w, https://substackcdn.com/image/fetch/$s_!KTzi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82d62528-81aa-4c34-8a48-e324ac35a66c_1887x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!KTzi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82d62528-81aa-4c34-8a48-e324ac35a66c_1887x1000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KTzi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82d62528-81aa-4c34-8a48-e324ac35a66c_1887x1000.png" width="1887" height="1000" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82d62528-81aa-4c34-8a48-e324ac35a66c_1887x1000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1000,&quot;width&quot;:1887,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!KTzi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82d62528-81aa-4c34-8a48-e324ac35a66c_1887x1000.png 424w, https://substackcdn.com/image/fetch/$s_!KTzi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82d62528-81aa-4c34-8a48-e324ac35a66c_1887x1000.png 848w, https://substackcdn.com/image/fetch/$s_!KTzi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82d62528-81aa-4c34-8a48-e324ac35a66c_1887x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!KTzi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82d62528-81aa-4c34-8a48-e324ac35a66c_1887x1000.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><p><strong>Why RAG, Not Just Upload to Claude</strong></p><p>The obvious approach is uploading the PDF to Claude or ChatGPT and asking questions. It works, mostly. But it has three problems.</p><p>First, the SpaceX S-1 was filed after most model training cutoffs. For specific figures the model has no training data &#8212; it either says "I don't know" or hallucinates a plausible number. I benchmarked this: asking Claude directly about SpaceX's 2025 revenue without context produces a confident wrong answer.</p><p>Second, a 395-page (after amendments) document strains context windows. Models start losing details from the middle of the document when they're trying to hold everything at once. Important disclosures on pages 80-200 get deprioritized for content near the beginning and end.</p><p>Third, citations are vague. "According to the filing" isn't useful when you're trying to verify a specific governance claim before an IPO.</p><p>RAG solves all three. You precompute the embeddings once, retrieve only the relevant chunks at query time, and the model sees focused context rather than 395 pages of noise.</p><div><hr></div><p><strong>Why Not Fine-Tune</strong></p><p>Before settling on RAG, I considered fine-tuning a smaller model on the filing content. The results from my own benchmarking &#8212; fine-tuning Mistral-7B on 25 SpaceX Q&amp;A pairs &#8212; ruled it out quickly.</p><p>Fine-tuning on a document teaches the model to reproduce facts it has seen during training. Ask it a question that maps closely to a training example and it answers well. Ask it anything slightly outside that distribution &#8212; a follow-up question, a cross-reference between sections, a question phrased differently &#8212; and it hallucinates confidently. The model has memorized, not understood.</p><p>RAG sidesteps this entirely. The model never sees the filing during training. At query time, relevant chunks are retrieved and injected as context. The model reads those chunks and answers from them. It's closer to open-book exam than memorization &#8212; and for a 395-page legal document with dense cross-references, open-book is the right approach.</p><p>Fine-tuning also has a practical problem for this use case: when SpaceX files an amendment &#8212; which they did twice within two weeks &#8212; the fine-tuned model is immediately stale. Re-ingesting a RAG pipeline takes under 5 minutes. Re-fine-tuning a model takes hours and compute budget.</p><div><hr></div><p><strong>The Architecture</strong></p><pre><code>SpaceX S-1 PDF (395 pages &#8594; 871 chunks)
    &#8595; pdfplumber &#8212; extract text page by page
    &#8595; sliding window chunker &#8212; 400 words, 100 overlap
    &#8595; all-MiniLM-L6-v2 &#8212; embed chunks &#8594; 384-dim vectors
    &#8595; Qdrant Cloud &#8212; store 871 vectors + page metadata

User question
    &#8595; all-MiniLM-L6-v2 &#8212; embed query
    &#8595; cosine similarity &#8594; top 15 candidates
    &#8595; re-rank &#8212; penalize summary pages
    &#8595; Claude Haiku &#8212; generate cited answer
    &#8595; &#177;8 page range citation</code></pre><p>Four components. Each does one thing.</p><p>Two separate models &#8212; intentional design. all-MiniLM-L6-v2 handles embeddings only. Claude Haiku handles generation only. Embedding models are optimized for semantic similarity &#8212; small, fast, deterministic, 384 dimensions. Generation models are optimized for instruction following and text quality. Using the same model for both would mean either a slow embedding step or a weak generation step. Keeping them separate is standard RAG practice and worth being explicit about.</p><p>Why Qdrant. Qdrant's free tier is generous enough for a single filing (871 chunks, 384 dimensions). The HNSW index makes similarity search fast at this scale. Local Qdrant works for development &#8212; Qdrant Cloud for production without managing infrastructure.</p><p>395 pages &#8594; 871 chunks. Average 2.2 chunks per page after the sliding window. Total vectors stored: 871 &#215; 384 dimensions. Each chunk stores text, page number, and end page in the payload &#8212; retrieved alongside the vector for citation generation.</p><div><hr></div><p><strong>The Chunking Decision</strong></p><p>400 words per chunk with 100-word overlap. Why these numbers?</p><p>Smaller chunks (200 words) lose context for multi-sentence financial disclosures. A revenue figure appears on one line; the explanation &#8212; segment breakdown, YoY comparison, key drivers &#8212; spans the next five sentences. Split at 200 words, you retrieve the number without the context.</p><p>Larger chunks (800 words) reduce retrieval precision. You retrieve more text than you need and dilute the relevant signal with adjacent content.</p><p>The 100-word overlap ensures no fact gets cut at a chunk boundary without appearing in an adjacent chunk. Any sentence that spans two chunks will be fully retrievable from either side.</p><div><hr></div><p><strong>Why Claude Haiku for Generation</strong></p><p>I benchmarked five LLMs on <strong>15 SpaceX S-1 questions</strong> spanning factual recall, multi-step reasoning, and structured output. Each model received the same RAG context, and I measured both answer quality and end-to-end latency.</p><p>Model                Score   Latency</p><p>------------------------------------</p><p>Claude Haiku         4.7/5    2.8 s</p><p>phi4:14b (local)     4.5/5   27.6 s</p><p>qwen2.5:14b (local)  4.4/5   26.9 s</p><p>mistral:7b (local)   4.4/5    9.0 s</p><p>deepseek-r1:14b      4.3/5  102.8 s</p><p>The quality gap between Haiku and local 14B models is 0.2 points. The latency gap is 10x. For a web product where users are waiting for an answer, Haiku wins decisively.</p><p>One interesting finding: structured output scores were nearly identical across all models (4.4-4.6). The differentiation came entirely from factual accuracy and reasoning &#8212; where Haiku's training data and instruction following consistently outperformed locally-run open models.</p><div><hr></div><p><strong>The Challenges</strong></p><p><strong>The summary pages problem.</strong></p><p>The executive summary (pages 1-24) mentions every major topic at a high level &#8212; consistently scoring highest in semantic similarity for almost any query, even when detailed content existed 100+ pages later.</p><p>Fix: retrieve 15 candidates, then apply a 0.15 penalty to chunks from pages under 25. Most substantive disclosures live deeper in the filing. Penalizing the summary section keeps retrieval focused on the narrative sections where specific claims and governance details actually appear.</p><p><strong>The page citation problem.</strong></p><p>The most challenging aspect was generating accurate page citations. The core issue: the SEC EDGAR filing only exists as HTML, which I converted to PDF using Chrome's print function. Chrome's HTML reflow during rendering means the text layer in the PDF doesn't always align with what you see visually.</p><p><strong>What I tried first &#8212; standalone number regex</strong></p><p>The first attempt looked for standalone numbers at the bottom of each page. Failed immediately &#8212; financial tables, footnote numbers, and reference counts appear throughout the page content including near the bottom. Too many false positives to be reliable.</p><p><strong>What I tried second &#8212; Chrome's </strong>N/313<strong> footer regex</strong></p><p>Chrome adds <code>N/313</code> page indicators in the footer during printing. I wrote a regex to extract it.</p><p>In theory this pattern is unique and can't appear elsewhere in the filing. In practice it was unreliable &#8212; the footer text wasn't always cleanly captured by pdfplumber's text extraction, so the regex frequently missed pages.</p><p><strong>What I tried third &#8212; WeasyPrint HTML&#8594;PDF conversion</strong></p><p>WeasyPrint converts HTML to properly paginated PDF where the text layer and visual layer are aligned by design. This would have eliminated the problem entirely. Failed on macOS &#8212; requires GTK libraries (<code>libgobject</code>, <code>pango</code>, <code>cairo</code>) that don't install cleanly on macOS without significant dependency management. Abandoned after an hour of dependency hell.</p><p><strong>What I tried fourth &#8212; paged.js</strong></p><p>A JavaScript library specifically designed for CSS-based HTML pagination. More macOS-friendly than WeasyPrint. The 11.8MB HTML filing with separately hosted image assets made this impractical &#8212; the converted PDF would be missing all images and the pagination would differ from the original rendering anyway.</p><p><strong>What actually works &#8212; position-based extraction</strong></p><p>The winning approach uses pdfplumber's coordinate system directly. Instead of parsing text, it looks for a standalone digit in the bottom 10% of the page, centered between 20&#8211;80% of the page width.</p><p>This reliably catches the printed page number without depending on text extraction of footer lines. Citations display a &#177;8 page range to account for any remaining rendering uncertainties.</p><p><strong>Demo card caching</strong></p><p>The /api/demo route is intentionally cached by Next.js. The three demo questions are fixed, the underlying data doesn't change between ingestion runs, and the answers are expensive to generate &#8212; hitting both Qdrant and the Claude API on every page load would add latency for no benefit. Cached results mean the landing page loads fast every time.</p><p><strong>The filing is a moving target.</strong></p><p>SpaceX filed two amendments after the original S-1 &#8212; S-1/A #1 on June 1 and S-1/A #2 on June 3 &#8212; with updated financials and the IPO price range ($135/share). The RAG pipeline re-ingests any filing version in under 5 minutes. When Anthropic and OpenAI file their S-1s later this year, the same pipeline handles them.</p><div><hr></div><p><strong>Conversation Memory</strong></p><p>The app maintains conversation history across turns. Follow-up questions work without re-explaining context &#8212; "which segment is most profitable?" after asking about revenue breakdown uses the prior exchange. History is passed as the Anthropic messages array, capped at the last 10 exchanges to keep context window usage bounded.</p><div><hr></div><p><strong>Stack</strong></p><ul><li><p><strong>Frontend:</strong> Next.js 14 on Railway. Migrated from a Streamlit prototype &#8212; easier to keep the same platform than migrate.</p></li><li><p><strong>Vector storage:</strong> Qdrant Cloud. Free tier covers a single filing comfortably. HNSW index, no infrastructure to manage.</p></li><li><p><strong>Generation:</strong> Anthropic API (Claude Haiku). Chosen on latency and quality benchmarks above.</p></li><li><p><strong>Embeddings:</strong> @xenova/transformers running all-MiniLM-L6-v2 in Node.js. Runs entirely locally &#8212; no embedding API calls at query time, which reduces latency and cost per query. Ingestion is separated from retrieval; embeddings are computed once and pushed to Qdrant Cloud.</p></li><li><p><strong>Domain:</strong> Cloudflare. AskS1.com at ~$10/year.</p></li></ul><div><hr></div><p><strong>What's Next</strong></p><p>Anthropic and OpenAI S-1s are expected soon. AskS1 will be there when they file.</p>]]></content:encoded></item></channel></rss>