Gizmo: Self-Awareness as Emergent Property
Most agents know what they can do because a developer told them. Gizmo discovers its own capabilities by searching. The difference is architectural.
I asked Gizmo what tools it had available. It did not recite a list. It searched.
The agent fired a NoteSearch query — the same query it uses to find user memories, reference documents, and session context — and discovered its own capabilities by reading the results. It found tool descriptions stored as notes, skill definitions stored as notes, and behavioral instructions stored as notes. The answer to “what can you do?” came from the same retrieval path as the answer to “what do you know about the user?”
That moment was not programmed. It emerged from an architecture decision I made months earlier: everything is a note.
The Architecture Decision
Gizmo NXT stores six types of knowledge — global instructions, session instructions, memories, chat extracts, skills, and documents — in a single Core Data entity called GNNote. Every piece is markdown with YAML frontmatter. The type is a metadata field, not an architectural boundary. Storage is identical. Retrieval is identical. Ranking is identical.
This sounds like a data modelling choice. It is a cognitive architecture choice.
In every other agent system I have studied, self-awareness is engineered explicitly. Letta pins persona blocks in the context window. LangChain agents receive their tool descriptions as part of the system prompt. OpenAI assistants get their instructions injected at the top of every conversation. The agent “knows” what it can do because the developer declared it at design time, in a static configuration that the agent cannot query, cannot reason about, and cannot update without a redeployment.
In Gizmo, the agent’s instructions are notes. Its skills are notes. Its behavioral guidelines are notes. And it discovers them the same way it discovers anything else — by searching its unified knowledge store through NoteSearch.
The Mechanism
NoteSearch runs a retrieval stack built entirely on Apple’s native NLP frameworks. NLTokenizer extracts semantic keywords from the query. NLTagger identifies named entities — people, places, projects, technical concepts. A four-factor reranking function scores every candidate note: semantic relevance at 40%, recency at 25%, frequency of access at 15%, title match at 20%. A maximum marginal relevance algorithm enforces diversity so the results do not cluster around a single angle.
No embeddings. No vector database. No RAG pipeline. No external API calls. The entire stack runs locally on Apple silicon, using frameworks that ship with every Mac and iPhone. The retrieval is fast because it is simple. It is accurate because the notes were written to be found — structured at write time, not guessed at query time.
The critical insight: when the agent searches for context and a skill note matches, the agent discovers its own competence in real time. When a behavioral instruction matches, the agent rediscovers its own operating principles. The self-model is not a fixed declaration loaded at boot. It is a living query result that evolves at the speed of note creation.
Why This Matters
Add a new skill to Gizmo and it is immediately discoverable. No redeployment. No configuration change. No restart. The agent’s next search might surface a capability it did not have an hour ago — and it will reason about that capability alongside its memories, its instructions, and its reference documents, because they all live in the same store and respond to the same query.
Remove a skill and it disappears from the agent’s self-model just as cleanly. The agent does not need to be told it lost a capability. The search simply stops returning it.
This is the difference between a capability manifest and a knowledge space. A manifest is a static document that drifts from reality the moment the system changes. A knowledge space is always current because the search always runs against the current state. The agent’s self-awareness tracks its actual capabilities — not what someone remembered to list in a configuration file three deployments ago.
The Cross-Domain Parallel
A guitarist who plays the same instrument for twenty years does not think about what the guitar can do. The knowledge is in the hands. Ask them to describe their technique and they will play something — they will search their muscle memory and discover the answer in real time, the same way they discover a melody or a chord voicing.
That is closer to what Gizmo does than what any static-prompt agent does. The agent does not carry a laminated card listing its capabilities. It plays, and in playing, discovers what it knows.
What the Industry Gets Wrong
Most agent architectures separate what the agent knows from what the agent is from what the agent can do. Three systems, three retrieval paths, three maintenance burdens. The fragmentation is not just an engineering inconvenience. It is a cognitive ceiling. An agent that cannot search its own capabilities the same way it searches its memories cannot reason about what it knows versus what it can do. It cannot connect a skill it learned last week to a question being asked right now.
Gizmo collapses that boundary. One entity. One store. One search. The agent does not switch between modes of cognition. It searches its unified knowledge and acts on what it finds.
Self-awareness was not a feature I designed. It was a consequence of refusing to build separate systems for separate kinds of knowledge. The architecture made the decision. The agent made it real.
The most interesting properties of a system are the ones you did not plan. You just built the conditions for them to emerge.
More from this domain
12 Aug 2026
MachinesThe Nine-Stage Mind
Nine stages around a 27B local model. On my own tasks it beats a trillion-parameter model alone. Architecture beats parameter count.
9 Jul 2026
MachinesThe Harness Is the Product
Every team has Claude Sonnet 4.5. A 20-step pipeline at 95% per step finishes 36% of the time. The harness is where the product lives.
11 Jun 2026
MachinesM3: Frontier Coding, Not Frontier Freedom
428B params, 1M context, 80.5% on SWE-Bench Verified by its own card. M3 wins long-context coding — then drops MIT for a license with strings.