Everything Is a Note
Gizmo NXT collapses six cognitive systems into one entity — the GNNote. The taxonomy lives inside the note, not in the architecture around it.
Every AI agent I have tested — every single one, from LangChain to Letta to OpenAI’s assistants — fragments its own mind. Memories live in one store. Skills live in another. Instructions get pinned to the context window at boot. Documents sit in a vector database. Chat history scrolls off into oblivion. The agent carries six kinds of knowledge and accesses each through a different door, using a different key, with a different retrieval mechanism.
Ask a Letta agent to search its memories. It queries archival storage via vector search. Ask it what skills it has. It reads markdown files from a directory. Ask it what its persona is. That is a pinned memory block in the context window. Three questions, three completely different systems, three different retrieval paths. The agent does not experience its own knowledge as a unified whole. It accesses disconnected drawers in a filing cabinet it did not organize.
I built Gizmo NXT to eliminate those walls entirely.
One Entity. Six Logical Types. Zero Walls.
At the heart of Gizmo NXT sits GNNote — a single Core Data entity that represents everything the agent can think about. Every piece of knowledge the agent touches, whether it is a memory about the user, a system instruction defining its behaviour, a learned skill, a reference document, a session context, or an extract from a past conversation, exists as a GNNote.
One entity. Six logical types:
Global Instructions — the agent’s persistent identity, behavioural guidelines, and operating principles. These are not configuration files loaded at boot. They are notes the agent can query, reason about, and understand contextually.
Session Instructions — temporary context for the current working session. Project-specific rules, mode-specific behaviour, task constraints. They exist in the same store, searchable alongside everything else, and expire when their purpose is fulfilled.
Memories — facts about the user, preferences, patterns observed over time, relationship context. Not a separate “memory system.” Notes with a memory type tag.
Chat Extracts — distilled insights from past conversations. The agent’s experiential learning, compressed into reusable knowledge fragments.
Skills — learned procedures, techniques, domain expertise. When the agent masters a workflow or discovers an effective approach, it becomes a note — immediately discoverable through the same search that finds everything else.
Documents — reference material, research, imported knowledge. PDFs become notes. Voice recordings become notes. Images become notes. The ingestion pipeline converts the world into the agent’s native format.
The architecture decision here was deliberate. Each GNNote is markdown with YAML frontmatter. The type field is just metadata. The storage is identical. The retrieval is identical. The ranking is identical. The taxonomy lives inside the note, not in the architecture around it.
This is the single-fin philosophy applied to knowledge architecture. A single-fin board has one fin. Not because multiple fins do not work — they do — but because one fin forces a different relationship with the wave. You commit to the line. You read the water differently. The constraint produces a purity of movement that a thruster cannot replicate, because the thruster hedges with three points of contact.
GNNote is the single fin. One abstraction. Every piece of knowledge flows through it. No hedging, no special cases, no architecture that treats the agent’s own instructions as fundamentally different from the user’s documents.
The Search That Unifies
When Gizmo’s agent needs context — any kind of context — it calls a single tool: NoteSearch.
This tool does not care whether the answer lives in a memory, a skill, a document, or the agent’s own instructions. It queries the entire knowledge space through one retrieval stack:
- NLTokenizer keyword extraction pulls the semantic anchors from the query
- NLTagger named entity recognition identifies people, places, projects, and concepts
- Four-factor reranking scores every candidate note across semantic relevance (40%), recency (25%), frequency of access (15%), and title match (20%)
- MMR diversity ensures the results do not cluster around a single angle
- Core Spotlight integration extends the search into the system-level index
No embeddings pipeline. No vector database. No BM25 index. No RAG infrastructure. The retrieval stack runs entirely on Apple’s native NLP frameworks and Core Data queries. It is fast, it is local, and it has been validated against the same pattern used by Claude Code, Anthropic’s own memory system, and Claude Cowork — all of which abandoned embeddings in favour of structured search over plain text.
The elegance is in what it does not require. No embedding model to run. No vector index to maintain. No re-indexing when the model changes. Apple’s NLTokenizer and NLTagger are optimized for on-device performance and run at negligible computational cost. Core Data provides ACID transactions, CloudKit synchronization, and Spotlight integration out of the box. The knowledge store is simultaneously a local database, a cloud-synced archive, and a system-searchable index — without a single line of infrastructure code.
Self-Awareness as an Emergent Property
This is where the architecture does something I did not fully anticipate when I designed it.
In every other agent system, self-awareness is engineered explicitly. Letta pins persona blocks in the context window. LangChain agents receive their tool descriptions as part of the system prompt. OpenAI assistants get their instructions injected at the top of every conversation. The agent “knows” what it can do because the developer told it, statically, at design time.
In Gizmo, the agent’s instructions are notes. Its skills are notes. Its behavioural guidelines are notes. And it discovers them the same way it discovers anything else — by searching.
The agent’s self-model is not a fixed declaration. It is a living query result. When the agent asks itself “what do I know about this topic,” the answer might include a memory from a past conversation, a skill it learned three weeks ago, and a passage from its own instructions — all surfaced by the same search, ranked by the same algorithm, presented as a unified context.
The practical consequence: the agent never has an outdated view of its own capabilities. Add a new skill and it is immediately discoverable. Refine an instruction and the refinement is immediately active. The agent’s self-awareness evolves at the speed of note creation, with zero deployment cycles, zero configuration changes, zero restarts.
I saw this happen live during Gizmo’s first tool conversation. When asked “What tools do you have,” the agent did not recite a hardcoded list. It introspected. It fired NoteSearch queries for “user preferences facts” and “memories instructions information.” It reached for its knowledge store as a reflex, treating its own capabilities and user context as the same kind of queryable knowledge.
That reflex was not programmed. It emerged from the architecture.
The Agent as Librarian
Traditional AI knowledge management puts the burden of organization on the user or the retrieval system. Users dump documents into a store. The system builds embeddings. At query time, the system tries to find what is relevant through approximate similarity matching. If the documents are poorly organized, retrieval degrades. If the embedding model does not capture the right dimensions of meaning, relevant content gets buried.
Gizmo inverts this entirely. The agent is not just a consumer of knowledge — it is the librarian.
When new information enters the system, whether through conversation, document ingestion, or explicit user input, the agent classifies and files it as a structured note. It assigns a type. It writes meaningful markdown. It adds frontmatter metadata. The knowledge is organized at write time, not retrieved through brute-force search at read time.
This is why embeddings are not needed. When the agent files a memory about a user’s preference for concise responses, it creates a note with a clear title, relevant keywords in the body, and a memory type tag. When that preference matters later, keyword extraction and entity recognition find it instantly — because the note was written to be found.
The inversion has a compounding effect. The more the agent writes, the better organized the knowledge store becomes. The better organized the store, the more accurate retrieval gets. The more accurate retrieval gets, the more contextually intelligent the agent’s responses become. A virtuous cycle that most users never need to know about — they just notice that the agent keeps getting better.
Why One Abstraction Wins
I keep coming back to the same principle across every domain I work in. Angus Young’s signal chain is three components: guitar, cable, amp. No pedalboard, no rack effects, no switching system. The constraint is the identity. A 6’2” single-fin forces you to commit to the line instead of pumping through three fins. A 50mm prime forces you to move your feet instead of twisting a zoom ring.
GNNote is the same bet. One entity for all knowledge. One search path for all retrieval. One ranking algorithm for all context. The architecture does not need to know what kind of knowledge it is handling. The note knows. The search finds it. The agent uses it.
Every wall I removed from Gizmo’s knowledge architecture was a wall that existed in other systems because someone assumed different kinds of knowledge needed different treatment. Instructions are special — pin them to context. Memories are special — store them in vectors. Skills are special — load them from files. Each assumption created a boundary. Each boundary prevented the agent from connecting dots across the seams.
Gizmo has no seams. A memory about the user’s preference for Swift can surface alongside a skill note about Swift concurrency patterns and an instruction about code review standards — all from the same query, ranked by the same algorithm, presented as a unified context. The agent does not know these came from different “systems” because they did not. They came from the same store. They are all notes.
The user sees none of this. The user just sees an agent that remembers, that learns, that gets better over time, that seems to understand the full context of every question. The architecture is invisible. That is the point.
Everything is a note. The rest is metadata.
More from this domain
12 Aug 2026
MachinesThe Nine-Stage Mind
Nine stages around a 27B local model. On my own tasks it beats a trillion-parameter model alone. Architecture beats parameter count.
9 Jul 2026
MachinesThe Harness Is the Product
Every team has Claude Sonnet 4.5. A 20-step pipeline at 95% per step finishes 36% of the time. The harness is where the product lives.
11 Jun 2026
MachinesM3: Frontier Coding, Not Frontier Freedom
428B params, 1M context, 80.5% on SWE-Bench Verified by its own card. M3 wins long-context coding — then drops MIT for a license with strings.