The Search That Sees Everything
Gizmo's NoteSearch queries the entire knowledge space through one path. No embeddings, no vector database, no RAG pipeline.
The retrieval stack for most AI agents looks like this: a vector database (Pinecone, Chroma, Weaviate, pick your poison), an embedding model to convert text into high-dimensional vectors, a chunking strategy to slice documents into digestible pieces, a similarity search to find the nearest neighbours, and a RAG pipeline to inject the results into the model’s context window. Five moving parts before the agent has read a single word.
Gizmo’s retrieval stack looks like this: NoteSearch.
One tool. One query path. Every piece of knowledge the agent has ever encountered — instructions, memories, skills, chat extracts, documents, session context — searched through the same mechanism, ranked by the same algorithm, returned as a unified result set.
I did not start here. I started where everyone starts, planning an embeddings pipeline. Then I looked at what Claude Code does for memory retrieval. Then I looked at how Anthropic built memory into their own systems. Then I looked at Claude Cowork. All three abandoned embeddings. All three use structured search over plain text. Three teams at Anthropic, working on different products, independently converged on the same conclusion: for personal knowledge retrieval at the scale an individual agent handles, embeddings are overhead that buys you nothing.
I took that as a strong signal and built accordingly.
How NoteSearch Works
The query arrives. NoteSearch processes it through four layers, each running natively on Apple frameworks:
Keyword extraction. NLTokenizer breaks the query into tokens, identifies the semantic anchors — the words that carry meaning versus the ones that carry grammar. A query like “what did Bertrand say about Swift concurrency in the Gizmo architecture” yields anchors: Bertrand, Swift, concurrency, Gizmo, architecture. The stopwords fall away. The search terms emerge.
Named entity recognition. NLTagger runs across the extracted tokens and identifies entities — people, places, organizations, technical concepts. “Bertrand” is tagged as a person. “Swift” as a technology. “Gizmo” as a product. This enriches the search by allowing the system to weight entity matches differently from keyword matches. A note titled “Gizmo Architecture Decisions” that mentions Swift concurrency will score higher than a note that happens to contain the word “swift” in a different context.
Core Data query construction. The keywords and entities are combined into a compound NSPredicate that searches across note titles, body content, tags, and type metadata. This is not a full-text search engine. It is a structured query against a relational store, which means it returns exact matches and does not hallucinate proximity. The note either contains the term or it does not.
Four-factor reranking. Every candidate note that passes the query filter is scored across four dimensions: semantic relevance at 40%, recency at 25%, frequency of access at 15%, and title match at 20%. These are not learned weights. They are designed parameters with explicit rationale — tunable by the developer, transparent in their effect, deterministic in their output.
Relevance dominates because the most important note is the one that best answers the query. Recency matters because knowledge decays — a note from yesterday about the user’s current project is more useful than a note from six months ago about a completed one. Frequency of access captures importance through behaviour — notes the agent retrieves often are notes that matter often. Title match rewards precision — a note whose title directly addresses the query is almost certainly the right note.
After scoring, MMR diversity enforcement ensures the final result set does not cluster around a single angle. If five notes score highly but all address the same subtopic, the algorithm promotes lower-scoring notes that cover different facets. The agent sees the landscape of its knowledge, not just the nearest neighbours.
What It Does Not Require
This is where the architecture earns its keep.
No embedding model. NoteSearch does not convert text into vectors. It does not need a 400MB model loaded into memory to compute similarity. It does not produce approximate results that degrade unpredictably when the query falls outside the embedding model’s training distribution.
No vector database. No Pinecone subscription. No Chroma instance. No Weaviate cluster. No index maintenance. No re-indexing when the embedding model changes. No cold-start latency while the index loads.
No chunking strategy. Embeddings require documents to be sliced into chunks — 512 tokens, 1024 tokens, overlapping or not, with metadata carried through or lost at the boundary. Every chunking choice is a tradeoff between context preservation and retrieval granularity. NoteSearch does not chunk. Each GNNote is a complete unit of knowledge, written by the agent to be self-contained. The note is the chunk. The agent wrote it that way on purpose.
No RAG pipeline. No retrieval-augmented generation chain to debug. No prompt template that stitches retrieved chunks into a context window with separator tokens and relevance scores. The agent calls NoteSearch, receives ranked notes, and reads them. The pipeline is: search, read, respond.
Core Spotlight integration extends the search beyond Gizmo’s own store and into the system-level index. Notes indexed through Core Data are automatically available to Spotlight, which means the user can find Gizmo’s knowledge from the macOS search bar, from Siri, from any app that queries the system index. The knowledge store is not locked inside the agent. It is a first-class citizen of the operating system.
The Signal in the Simplicity
I keep returning to a principle that shows up in every domain I work in. The best signal chains are the shortest ones. Angus Young’s guitar tone comes from three components. Lindbergh’s portraits come from available light and no retouching. A single-fin rides on one point of contact with the water. The fewer things between the input and the output, the cleaner the signal.
NoteSearch is a short signal chain. Query comes in. Keywords and entities are extracted. Core Data returns matches. Four-factor scoring ranks them. Diversity enforcement balances the set. Notes come out. Six steps, all running locally, all deterministic, all auditable. I can inspect any search result and trace exactly why it was ranked where it was — which factor contributed what percentage, which keywords matched, which entity tags fired. There is no black box. Every cognitive step leaves a paper trail.
The Python AI ecosystem would have me running an embedding model, maintaining a vector index, tuning chunking parameters, optimizing retrieval prompts, and debugging a RAG chain — all to accomplish the same task that NoteSearch handles with Apple’s native NLP frameworks and a Core Data query. The overhead is not just computational. It is cognitive. Every additional component is a component that can fail, that needs monitoring, that requires expertise to tune.
I chose the path with fewer components. Not because I could not build the complex version. Because the simple version works better for what Gizmo actually does — searching one person’s knowledge, at the scale of thousands of notes, with sub-second response times, entirely on-device.
The search sees everything because everything is a note, and notes are what the search was built to find. No special cases. No separate retrieval paths. No architecture that treats different kinds of knowledge as different kinds of problem. One path. One algorithm. Full coverage.
That is the whole point. The elegance is not in what it does. It is in what it refuses to need.
More from this domain
12 Aug 2026
MachinesThe Nine-Stage Mind
Nine stages around a 27B local model. On my own tasks it beats a trillion-parameter model alone. Architecture beats parameter count.
9 Jul 2026
MachinesThe Harness Is the Product
Every team has Claude Sonnet 4.5. A 20-step pipeline at 95% per step finishes 36% of the time. The harness is where the product lives.
11 Jun 2026
MachinesM3: Frontier Coding, Not Frontier Freedom
428B params, 1M context, 80.5% on SWE-Bench Verified by its own card. M3 wins long-context coding — then drops MIT for a license with strings.