| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
Personal continually-learning local LLM. Knowledge in weights, not in prompts.
Early-stage research project. The architectural thesis — hybrid Mamba+attention base with three coordinated memory channels — works end-to-end on a MacBook (see empirical findings). Agentic / tool-using flows are an open problem.
Three things make cognit different from a frontier coding agent:
| Frontier agents | Cognit |
|---|---|
| Personalization in prompts (RAG, memory injection, system msgs) | Personalization in weights & state (LoRA adapter + persistent SSM state) |
| Memory grows your context cost on every call | Memory grows your adapter file (~1 MB), free at inference |
| Your data sits in prompts the vendor sees | Your data is baked into a local adapter you own |
| Personalization stops working when context fills | Adapter capacity scales with parameters, not context |
Cognit runs on a MacBook today. Hybrid Mamba+attention base (Zamba2-1.2B-Instruct-v2 by default), persistent per-session SSM state, LoRA fine-tuning between sessions, catastrophic-forgetting protection by default, OpenAI-compatible HTTP server for coding-agent integration.
pip install -e ".[server]"(Python 3.10+, PyTorch 2.2+, an Apple Silicon or CUDA GPU recommended. CPU works but training is slow.)
# First run downloads Zamba2-1.2B-Instruct-v2 (~2.4 GB) and creates the
# default profile. Subsequent runs reuse the cached model.
cognit init --yes
# Chat with it
cognit chat
# One-shot
cognit generate "Once upon a time"In the chat REPL, mark turns to teach cognit:
> Explain how our auth flow works cognit: ... > /good # mark the response as worth learning from > /fix It actually uses JWT, not OAuth. # correct the response > /exit # session ends; adapter trains on marked turns
Next session, cognit has absorbed those corrections into its weights.
Cognit exposes an OpenAI-compatible HTTP endpoint. Any agent that takes OPENAI_API_BASE works against it unchanged.
cognit serve(Default port is 11435 — one above Ollama's 11434, so the two can run side-by-side. Override with --port N if needed.)
Then in your agent's config:
export OPENAI_API_BASE=http://localhost:11435/v1
export OPENAI_API_KEY=not-needed
# model name: "cognit/default" (or any named profile)Or with the OpenAI Python SDK directly:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11435/v1", api_key="anything")
stream = client.chat.completions.create(
model="cognit/default",
messages=[{"role": "user", "content": "Refactor this function..."}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)Standard endpoints (OpenAI-compatible): /v1/chat/completions (with SSE streaming), /v1/models.
Cognit-specific extensions for learning:
Pi is a terminal coding agent from earendil-works that reads provider config from ~/.pi/agent/models.json. Cognit-on-Pi is an experimental integration right now: tool calling isn't implemented (see open problems), so Pi works well for paste-and-discuss flows (explaining code you paste, planning, conceptual coding questions) but not for autonomous file exploration. The launcher below is the easiest way to point Pi at cognit; on the first tool-using request cognit returns a one-time notice explaining the limit, then falls back to chat on subsequent requests.
Install Pi:
npm install -g @earendil-works/pi-coding-agent
# or: curl -fsSL https://pi.dev/install.sh | shThen point Pi at cognit with a single command:
cognit launch piThis starts cognit serve on port 11435 in the background (logs to ~/.cognit/launch-pi.log, with automatic fallback to a free port if something else holds 11435), registers cognit as a Pi provider while preserving any existing entries, sets it as Pi's default, and spawns pi interactively. On exit, the background server is stopped.
Flags:
If you'd rather wire it up by hand — e.g. to register multiple cognit profiles, or to keep Pi under your own config management — run cognit serve in one terminal, then add cognit to ~/.pi/agent/models.json:
{
"providers": {
"cognit": {
"baseUrl": "http://localhost:11435/v1",
"api": "openai-completions",
"apiKey": "cognit",
"models": [{ "id": "cognit/default" }]
}
}
}Merge with any existing providers entries (don't replace the whole file). To make cognit Pi's default, add to ~/.pi/agent/settings.json:
{
"defaultProvider": "cognit",
"defaultModel": "cognit/default"
}Then pi will route through cognit.
Cognit's memory isn't one mechanism — it's three, each operating at a different timescale and storing something different. The hybrid base isn't a stylistic choice; it's what makes this work at MacBook scale.
| Channel | Mechanism | What it stores | Persists across | Job |
|---|---|---|---|---|
| Within turn | Attention KV cache | Verbatim tokens, exact positions | Single forward pass | Exact recall — "what did the user just type 200 tokens back?" |
| Within session | SSM hidden state | Compressed running summary (gist) | Save/load — serialized to disk | Cheap long threads — resume tomorrow with conversation state intact |
| Across all sessions | LoRA adapter weights | Durable learned patterns | All future sessions, across reboots | Knowledge in weights — the model gets sharper at you over weeks |
These are not interchangeable. Don't try to smuggle exact-recall into the SSM channel; don't expect LoRA to carry the gist of an ongoing conversation. Each layer has its job.
Each channel above depends on a property of the base architecture:
This is why cognit's session save/load works and why the whole system fits on a MacBook: SSM state is what serializes, hybrid runtime costs stay bounded as threads grow, and a 1-3B hybrid base plus a small LoRA can do inference and online learning in laptop-class memory.
SSM state across sessions could in principle carry knowledge forward. In practice it doesn't, for the same reason gist-summaries don't substitute for studying: it's a lossy compression, unstructured, and the next conversation will overwrite it. What you actually want preserved across weeks of use — your codebase quirks, your writing style, terminology you've corrected — needs to live in weights. LoRA is that channel.
The base model stays frozen; only a small rank-r adapter (~1 MB at Zamba2-1.2B) learns. Adapter updates are cheap enough to run as background work after a chat session ends.
When a turn is promoted to the capture queue (you /good or /fix it in chat, or an agent calls POST /mark), cognit runs a few gradient steps of LoRA fine-tuning on the accumulated captures — on session exit, or via explicit train_pending. The base model's weights are frozen; only the small adapter changes.
The same training loop has two feeders, both writing into the same LoRA:
Conversation capture handles correction ("you got this wrong, here's right"); corpus ingestion handles bootstrapping ("here's everything I want you to already know on day one"). Both land in the same adapter and coexist seamlessly.
qmd as the backend — rather than a hand-rolled markdown walker — gives cognit stable qmd://collection/path URIs, dedup, and shared state with anything else you've connected to qmd (search, MCP, retrieval). See examples/06_growing_brain_demo.py for the full loop end-to-end.
Adapter interference is the real engineering risk: earlier-learned patterns in the LoRA can be overwritten by later updates. Cognit's mitigation is replay sampling — mix previously-trained captures into each new training pass — which is trivial (the captures are already in the queue) but ~99% effective in our measurement on Zamba2-1.2B vs unprotected. L2 regularization toward pre-pass weights was nearly useless at typical LoRA scales (kept as an opt-in research knob, not the default). See examples/08_forgetting_protection_demo.py.
The three-channel thesis isn't just architectural framing — both non-trivial channels (SSM, LoRA) have been demonstrated end-to-end on Zamba2-1.2B running locally. Reproduce with the example scripts; numbers below are from M-series MPS runs.
examples/04_persistence_demo.py
Pre-fill a ~500-character prior context into a session, persist the SSM state to disk on exit, then load it back into a new session. Generate from a neutral prompt with and without the loaded state. Findings:
This validates the SSM channel as actual working memory across save/load. A caveat baked into Zamba2's hybrid design: attention-heavy hybrid layers make SSM-only persistence a "gist" memory rather than exact recall — by design (see the architecture section above).
examples/05_lora_learning_demo.py
Baseline: generate from "The Glorpfish lives in" — Zamba2 confabulates something (the Glorpfish is fictional). Then teach cognit 3 corrected turns about it; session exit triggers a few gradient steps of LoRA fine-tuning. Open a fresh session, generate from the same prompt:
This validates the LoRA channel as durable across-session learning. A small artifact, layered on the shared frozen base, encodes what's been taught.
Each profile (work, personal, research, …) has its own adapter, captures, and sessions. Switching profiles is like switching to a different model — no cross-contamination.
cognit profile new work --model-type zamba2 --base Zyphra/Zamba2-1.2B-Instruct-v2
cognit profile new personal --model-type zamba2 --base Zyphra/Zamba2-1.2B-Instruct-v2
cognit --profile work chat --session pr-review
cognit --profile personal generate "What do I usually cook on Tuesdays?"Within a profile, sessions are conversation threads — separate SSM state per thread, shared LoRA. Layout:
~/.cognit/profiles/<name>/
├── profile.json # config
├── adapter.pt # LoRA weights for this profile
├── captures.jsonl # capture queue
└── sessions/
├── default.safetensors # SSM state per thread
└── pr-review.safetensors
| Script | What it shows |
|---|---|
| 04_persistence_demo.py | SSM state persistence empirically changes generation |
| 05_lora_learning_demo.py | LoRA training durably encodes new knowledge |
| 06_growing_brain_demo.py | Corpus ingestion + conversation capture → same adapter |
| 07_profile_isolation_demo.py | Profile isolation (work vs personal) |
| 08_forgetting_protection_demo.py | Catastrophic-forgetting protection measured 4 ways |
All examples run against Zamba2-1.2B (no checkpoint setup beyond cognit init --yes).
Alpha. The system works end-to-end on a MacBook with Zamba2-1.2B today. What's known:
What's missing or rough:
The architectural thesis is validated end-to-end (see empirical findings). What this project hasn't yet answered:
Cognit wraps a frozen hybrid Mamba+attention base model with three wrapper layers:
At inference, the frozen base + your profile's LoRA produce the response. The session's attention KV builds within the forward pass and discards after; the SSM state updates in place and is checkpointed to disk on session save. Marked turns feed the capture queue; on session exit, captures train the adapter (with replay protection). Next session starts with the updated adapter and the previous SSM state restored.
MIT — see LICENSE.
| Back | FazBrowse Home | New Git URL |