An AI agent that uses Sakana.ai Doc-to-LoRA to embed document knowledge into model weights instead of consuming context. I also trained a new Doc-to-LoRA hypernetwork for Ministral-3-3B-Instruct-2512 from scratch and reimplemented the inference pipeline in MLX for native Apple Silicon support.
LLMs face constant context size pressure. As conversations grow, agentic systems must compact or summarize prior context to stay within token limits, losing information along the way. Knowledge-augmented approaches like RAG make this worse by stuffing retrieved documents into prompts, consuming even more of the limited context window.
Three contributions: I ported Sakana AI's Doc-to-LoRA framework to Mistral's Ministral-3-3B-Instruct-2512, training a new hypernetwork from scratch with limited hackathon resources. I reimplemented the full inference pipeline for Apple Silicon (MLX), shipped as a standalone library. And I built Thoth, an agentic system that uses Doc-to-LoRA as a memory layer—powered by Mistral-7B and Sakana AI's pretrained hypernetwork.
A four-stage pipeline transforms raw documents into rank-8 LoRA adapters targeting the model's MLP layers. The same architecture applies to both my Ministral-3-3B-Instruct-2512 training and the Mistral-7B hypernetwork used in Thoth.
Document text is tokenized using the target model's tokenizer
A frozen base model extracts per-layer hidden state activations
A Perceiver Resampler compresses tokens into 8 latent vectors
HyperLoRA produces LoRA A/B weight matrices for each layer
Doc-to-LoRA Architecture
+----------+ +------------------+ +--------------+ +-----------+
| Tokenize | -> | Base Model | -> | Perceiver | -> | HyperLoRA |
| | | (frozen encoder) | | Resampler | | |
+----------+ +------------------+ +--------------+ +-----------+
Document in, LoRA adapter out. Knowledge lives in weights, context stays free.
Thoth is a DSPy-based ReAct agent that uses Doc-to-LoRA as a memory layer. Instead of pasting documents into context, it converts them to LoRA adapters on-the-fly and queries the augmented model. For production reliability, Thoth uses Mistral-7B with Sakana AI's pretrained hypernetwork, running locally on Apple Silicon via MLX-LM.
Converts documents to LoRA adapters via Sakana AI's pretrained hypernetwork and loads them into the local Mistral-7B instance.
Queries the LoRA-augmented local model to retrieve knowledge stored in weights.
Falls back to DuckDuckGo web search when knowledge isn't stored in memory.
Fetches and cleans web page content, which can then be added to memory.
Thoth Agent Architecture
User Query
|
v
+-----------------------+ +---------------------+
| Mistral Large API | ---> | DuckDuckGo Search |
| (ReAct reasoning) | | Web Fetch |
| DSPy orchestration | +---------------------+
+-----------------------+
|
v add_memory / query_memory
+-----------------------+ +---------------------+
| Doc-to-LoRA | ---> | MLX-LM Server |
| Sakana AI pretrained | | Mistral-7B + LoRAs |
| hypernetwork | | (Apple Silicon) |
+-----------------------+ +---------------------+
As part of the hackathon, I reimplemented the entire Doc-to-LoRA inference pipeline for Apple Silicon, shipped as a standalone library inside Thoth at thoth/d2l. This enables the full hypernetwork pipeline to run natively on Mac without CUDA, powering Thoth's local-first architecture.
Full reimplementation of the Perceiver Resampler (9 encoder blocks, GQA attention, GatedMLP) and the HyperLoRA network that generates rank-8 LoRA A/B weight matrices for all 32 transformer layers.
Extracts per-layer hidden states from a frozen Mistral-7B, tokenizing documents with the model's chat template and collecting activations across all 32 layers.
Converts generated LoRA weights to mlx-lm compatible format—SafeTensors binaries with JSON config—ready to be hot-loaded into the running MLX-LM server.
Downloads and loads Sakana AI's pretrained hypernetwork from Hugging Face, handling weight format conversion and non-strict state dict loading.
| Aspect | Doc-to-LoRA | RAG |
|---|---|---|
| Context Window | Free — knowledge in weights | Consumed by retrieved chunks |
| Latency | Sub-second generation + normal inference | Retrieval + reranking per query |
| Knowledge Depth | Full document absorbed | Limited to retrieved snippets |
| Composability | Multiple LoRAs compose along rank | Limited by context budget |
| Best For | Repeatedly-queried documents | Broad, diverse knowledge bases |
Separately from Thoth, I ported the Doc-to-LoRA framework to Ministral-3-3B-Instruct-2512, training a new hypernetwork from scratch. The training uses context distillation: a teacher model reads documents and answers questions, then the student (hypernetwork-generated LoRA, no document in context) tries to match the teacher's output distribution via KL divergence loss. While Thoth uses Sakana AI's mature Mistral-7B hypernetwork.
Self-generated QA pairs via Ministral-3-3B-Instruct-2512 (vLLM) on four datasets: SQuAD (Wikipedia QA), DROP (discrete reasoning), ROPES (science cause-effect), PwC (academic papers).
4× NVIDIA A100 80GB GPUs. Gradient accumulation of 8, max context length 2,048 tokens, L1 regularization 0.1. Tracked with Weights & Biases.
Adapting the original Sakana AI codebase to Mistral's architecture required solving several non-trivial challenges.
Ministral-3-3B-Instruct-2512 ships as a multimodal Pixtral variant. I extracted the text-only MistralForCausalLM component for use as the encoder.
Default FP8-quantized checkpoint produced incoherent outputs. Switched to the official BF16 variant with a model-mapping system.
Tekken v13 tokenizer required upgrading mistral-common to 1.9.0+ and patching vLLM assertion checks.
transformers 4.51.3 doesn't recognize "ministral3"—registered it as MistralConfig at import time.
Hypernetwork training code. Ports Sakana AI's Doc-to-LoRA to Ministral-3-3B-Instruct-2512 with full training pipeline and configs.
DSPy ReAct agent using Doc-to-LoRA as a memory layer. Uses Mistral-7B + Sakana AI's pretrained hypernetwork, running locally on Apple Silicon via MLX-LM.
Trained ~309M parameter Perceiver hypernetwork for Ministral-3-3B-Instruct-2512. Hackathon research contribution, separate from the Sakana AI hypernetwork used in Thoth.
Full training metrics, loss curves, and hyperparameter details for the 4,000-step training run.
"Doc-to-LoRA: Sub-Second Knowledge Injection into LLMs via Document-to-LoRA Translation" — Sakana AI, February 2026.