Mistral AI Worldwide Hackathon 2026

Thoth & mlx-D2L

An AI agent that uses Sakana.ai Doc-to-LoRA to embed document knowledge into model weights instead of consuming context. I also trained a new Doc-to-LoRA hypernetwork for Ministral-3-3B-Instruct-2512 from scratch and reimplemented the inference pipeline in MLX for native Apple Silicon support.

The Problem & The Solution

! The Problem

LLMs face constant context size pressure. As conversations grow, agentic systems must compact or summarize prior context to stay within token limits, losing information along the way. Knowledge-augmented approaches like RAG make this worse by stuffing retrieved documents into prompts, consuming even more of the limited context window.

✓ My Solution

Three contributions: I ported Sakana AI's Doc-to-LoRA framework to Mistral's Ministral-3-3B-Instruct-2512, training a new hypernetwork from scratch with limited hackathon resources. I reimplemented the full inference pipeline for Apple Silicon (MLX), shipped as a standalone library. And I built Thoth, an agentic system that uses Doc-to-LoRA as a memory layer—powered by Mistral-7B and Sakana AI's pretrained hypernetwork.

Doc-to-LoRA Pipeline

A four-stage pipeline transforms raw documents into rank-8 LoRA adapters targeting the model's MLP layers. The same architecture applies to both my Ministral-3-3B-Instruct-2512 training and the Mistral-7B hypernetwork used in Thoth.

1

Tokenize

Document text is tokenized using the target model's tokenizer

2

Encode

A frozen base model extracts per-layer hidden state activations

3

Compress

A Perceiver Resampler compresses tokens into 8 latent vectors

4

Generate

HyperLoRA produces LoRA A/B weight matrices for each layer

                          Doc-to-LoRA Architecture

  +----------+    +------------------+    +--------------+    +-----------+
  | Tokenize | -> |    Base Model    | -> |   Perceiver  | -> | HyperLoRA |
  |          |    | (frozen encoder) |    |   Resampler  |    |           |
  +----------+    +------------------+    +--------------+    +-----------+

  Document in, LoRA adapter out. Knowledge lives in weights, context stays free.

Thoth: LoRA-Memory Agent

Thoth is a DSPy-based ReAct agent that uses Doc-to-LoRA as a memory layer. Instead of pasting documents into context, it converts them to LoRA adapters on-the-fly and queries the augmented model. For production reliability, Thoth uses Mistral-7B with Sakana AI's pretrained hypernetwork, running locally on Apple Silicon via MLX-LM.

✍ add_memory

Converts documents to LoRA adapters via Sakana AI's pretrained hypernetwork and loads them into the local Mistral-7B instance.

🔍 query_memory

Queries the LoRA-augmented local model to retrieve knowledge stored in weights.

🌐 web_search

Falls back to DuckDuckGo web search when knowledge isn't stored in memory.

📄 web_fetch

Fetches and cleans web page content, which can then be added to memory.

                        Thoth Agent Architecture

  User Query
      |
      v
  +-----------------------+      +---------------------+
  |  Mistral Large API    | ---> |  DuckDuckGo Search  |
  |  (ReAct reasoning)    |      |  Web Fetch          |
  |  DSPy orchestration   |      +---------------------+
  +-----------------------+
              |
              v  add_memory / query_memory
  +-----------------------+      +---------------------+
  |  Doc-to-LoRA          | ---> |  MLX-LM Server      |
  |  Sakana AI pretrained |      |  Mistral-7B + LoRAs |
  |  hypernetwork         |      |  (Apple Silicon)    |
  +-----------------------+      +---------------------+

MLX Doc-to-LoRA Library

As part of the hackathon, I reimplemented the entire Doc-to-LoRA inference pipeline for Apple Silicon, shipped as a standalone library inside Thoth at thoth/d2l. This enables the full hypernetwork pipeline to run natively on Mac without CUDA, powering Thoth's local-first architecture.

Perceiver & HyperLoRA

Full reimplementation of the Perceiver Resampler (9 encoder blocks, GQA attention, GatedMLP) and the HyperLoRA network that generates rank-8 LoRA A/B weight matrices for all 32 transformer layers.

Context Encoder

Extracts per-layer hidden states from a frozen Mistral-7B, tokenizing documents with the model's chat template and collecting activations across all 32 layers.

MLX-LM Adapter Export

Converts generated LoRA weights to mlx-lm compatible format—SafeTensors binaries with JSON config—ready to be hot-loaded into the running MLX-LM server.

Checkpoint Loading

Downloads and loads Sakana AI's pretrained hypernetwork from Hugging Face, handling weight format conversion and non-strict state dict loading.

Doc-to-LoRA vs RAG

Aspect Doc-to-LoRA RAG
Context Window Free — knowledge in weights Consumed by retrieved chunks
Latency Sub-second generation + normal inference Retrieval + reranking per query
Knowledge Depth Full document absorbed Limited to retrieved snippets
Composability Multiple LoRAs compose along rank Limited by context budget
Best For Repeatedly-queried documents Broad, diverse knowledge bases

Ministral-3-3B-Instruct-2512 Hypernetwork Training

Separately from Thoth, I ported the Doc-to-LoRA framework to Ministral-3-3B-Instruct-2512, training a new hypernetwork from scratch. The training uses context distillation: a teacher model reads documents and answers questions, then the student (hypernetwork-generated LoRA, no document in context) tries to match the teacher's output distribution via KL divergence loss. While Thoth uses Sakana AI's mature Mistral-7B hypernetwork.

0.744
Final Train Loss
0.824
Final KL Loss
4,000
Training Steps
~8.2h
Training Time

Training Data

Self-generated QA pairs via Ministral-3-3B-Instruct-2512 (vLLM) on four datasets: SQuAD (Wikipedia QA), DROP (discrete reasoning), ROPES (science cause-effect), PwC (academic papers).

Infrastructure

4× NVIDIA A100 80GB GPUs. Gradient accumulation of 8, max context length 2,048 tokens, L1 regularization 0.1. Tracked with Weights & Biases.

Porting Gemma-2-2B → Ministral-3-3B-Instruct-2512

Adapting the original Sakana AI codebase to Mistral's architecture required solving several non-trivial challenges.

⚙ Multimodal Packaging

Ministral-3-3B-Instruct-2512 ships as a multimodal Pixtral variant. I extracted the text-only MistralForCausalLM component for use as the encoder.

⚡ FP8 Quantization Bug

Default FP8-quantized checkpoint produced incoherent outputs. Switched to the official BF16 variant with a model-mapping system.

🔧 Tokenizer Constraints

Tekken v13 tokenizer required upgrading mistral-common to 1.9.0+ and patching vLLM assertion checks.

📋 Config Registration

transformers 4.51.3 doesn't recognize "ministral3"—registered it as MistralConfig at import time.

Project Repositories

Mistral Large Mistral-7B Ministral-3-3B-Instruct-2512 DSPy MLX-LM vLLM PEFT / LoRA Perceiver PyTorch Weights & Biases Hugging Face

Based On

"Doc-to-LoRA: Sub-Second Knowledge Injection into LLMs via Document-to-LoRA Translation" — Sakana AI, February 2026.