Estimated reading time at 200 wpm: 5 minutes
I’m not going into why Authors may want to use AI. It’s big business out there. And no – it’s not about letting AI write books! FFS. But wait – it’s not only authors who may be interested in running an AI locally on their computers. AI can present tools for many use case scenarios. It’s all FREE. Right if that’s too good to be true, leave it alone. Go away and pay through the nose to the likes of ChatGPT. Well, I know the spiel, “I’m not an IT person. As for AI they just generate rubbish; never helped me.” If you’re one of those people, kindly depart and save yourself time for Instagram or Netflix, or wha’everrrr!
Whether or not you agree our Fat Disclaimer applies
In none of the following, I don’t need to know what each thing is. Most people who drive a car don’t need to understand all the electronics and parts of the engine that make it work. But may be you don’t drive a car, and won’t ever until you understand everything about how they work. That’s fine – you have individual choice! Like to leave now! This post is not meant to be helpful. It’s my notes. I put it here so I can find it in a flash if/when I need it in the future. If it helps one person, fine. If it helps nobody else, fine.
Running a local Large Language Model (LLM) for creative writing presents a unique challenge: balancing “model intelligence” against “memory capacity.” For an author, an AI that is brilliant but forgets the plot after three pages is useless.
This post details the technical process of optimizing a modern 12B parameter model for a system with 16GB RAM and a 10GB VRAM GPU. Do I know what all that means or understand it? I don’t. With AI assistance, the following was the work done. Yes – it takes time. What you want everything done instantly? The time spent will save me loads of time and headache as I expand my authorship career.
Caution: For idiots, AI software needs powerful computers and large storage. Their footprint can be 30 to 100GB. So not this stuff is not for ‘everybody’.
1. The Legacy Fallacy: Why Timestamps Matter
The initial candidate was a legacy quantisation of Mixtral 8x7B. While Mixtral was a landmark model, the specific files were tagged as “782 days old.” In the 2026 AI ecosystem, this represents a significant technical debt:
- Architecture Gaps: Older GGUF files often lack support for modern features like Flash Attention or specialised KV cache quantisation.
- Size vs. Resource Mismatch: The model weights (26.4 GB) far exceeded the available 10GB VRAM + 16GB RAM overhead, guaranteeing a “performance cliff” where the system would swap to disk, slowing generation to a crawl.
2. Model Selection: Finding the “Sweet Spot”
The goal was to find a model that maximises “literary reasoning” while fitting into a 10GB VRAM envelope.
- Selected Model:
Mistral-Nemo-12B-Instruct-v1 - Quantisation:
Q4_K_S(~7.12 GB) - Rationale: The 12B parameter count offers a significant upgrade in nuance over 7B or 8B models but is small enough to leave “room for memory” on a 10GB card. The
Q4_K_Sformat provides a 4-bit quantisation that preserves high accuracy while keeping the base weight under 7.5 GB.
3. Hardware Triage: Reclaiming VRAM
A common pitfall is ignoring background VRAM usage. The test system initially showed 8.9 GB / 10 GB utilised before loading the model, likely due to open browsers or GPU-accelerated applications. Closing these reclaimed roughly 6 GB of VRAM, which was the difference between a successful load and an immediate crash.
4. The “Long-Term Memory” Tweak: KV Cache Quantisation
This is the most critical setting for authors. By default, the AI stores the “history” of the conversation (the KV Cache) at high precision (FP16).
- The Tweak: Enabling 4-bit (Q4_0) KV Cache Quantisation.
- Technical Result: This compresses the memory cost of every token by 75%.
- Practical Impact: On a 10GB card, this allowed the Context Length to be safely increased from a measly 4,096 tokens (approx. 3,000 words) to 32,768 tokens (approx. 25,000 words). The AI can now “see” several chapters back into the story without exceeding the VRAM limit.
5. Software Configuration Summary
To replicate this author-centric environment in LM Studio (v0.4.0+), the following settings were applied:
| Setting | Value | Purpose |
|---|---|---|
| GPU Offload | Max (40 Layers) | Ensures the entire model runs on the fast graphics card. |
| Flash Attention | ON | Reduces memory overhead during long-context processing. |
| Context Length | 32,768 | Provides enough memory for character and plot tracking. |
| Temperature | 0.8 | Encourages creative, varied prose over rigid assistant-like speech. |
| System Prompt | Novelist Persona | Overrides the default “Helpful Assistant” to focus on sensory details and “Show, Don’t Tell” methodology. |
6. Monitoring the “VRAM Wall”
The finalised setup utilised 9.3 GB / 10.0 GB of VRAM. While this is “full,” the 4-bit KV cache ensures that generation speed remains high as the context window fills. If the usage hits 10.0 GB, the system is designed to spill the oldest tokens, maintaining stability for the user.
Technical Note to Self: Always check for newer model iterations (e.g., Mistral NeMo v2 or Llama 4 variants) as they become available, as the 700+ day gap in legacy models is usually indicative of obsolete quantisation logic.
Conclusion
Success! The model works fine. Now I can do lots of work offline and present a more distilled version for online AIs. That way I don’t burn through tokens online and get stopped out or invited to spend more money.
Others may find other tools for free online via LM Studio, for their own use case scenarios e.g. software development, graphic design etc.











