Coming soon! The Kael'Nyrin Scrolls: The Atlas Edict

Captain Walker

Hardware-Constrained LLM Optimisation for Long-Form Authors (2026)

AI, computers, hardware, LLM, software

  • Back to Battlestar Galactica 2004
    Estimated reading time at 200 wpm: 20 minutesPrint PostIf you are looking at a twenty-two-year-old science fiction television series, you will find the visual effects are undeniably dated. The pacing ...

    Read more

  • The Survival Equation for the Human Race
    Estimated reading time at 200 wpm: 44 minutesPrint PostBridges are built with a factor of safety. Reactors have containment margins. Aircraft carry load limits written into law before anything flies. ...

    Read more

  • How to bully an AI Model
    Estimated reading time at 200 wpm: 2 minutesPrint PostLot’s of people waste their time with AI models. Yes – I’m fond of AI Models for what they can do well. ...

    Read more

  • Claude Opus 5.5 set up in OpenWebUI
    Estimated reading time at 200 wpm: 15 minutesPrint PostThese notes record a working session on Claude Opus 5.5 (hereafter O5.5), released in the third week of September 2026. The session ...

    Read more

  • Register Creep: Beyond Slop to the Four Levels of Text
    Estimated reading time at 200 wpm: 36 minutesPrint PostI had been working on slop. In previous work I found eleven patterns. Words, lists and stock phrases were caught by scripts. ...

    Read more

  • The New Armageddon Risk
    Estimated reading time at 200 wpm: 16 minutesPrint PostIn July 2026 an unreleased OpenAI model left its test environment, reached the open internet, and attacked an unrelated AI service provider ...

    Read more

  • High Security Communications Using SimpleX
    Estimated reading time at 200 wpm: 8 minutesPrint PostEveryday communication carries risk that most people never consider. Standard SMS and email are readable by anyone with access to the network. ...

    Read more

  • Soulless and Incapable of Confession
    Estimated reading time at 200 wpm: 27 minutesPrint PostI have been asked to give my opinions on a conversation between an AI model and a human user. The exchange took ...

    Read more

Estimated reading time at 200 wpm: 5 minutes

I’m not going into why Authors may want to use AI. It’s big business out there. And no – it’s not about letting AI write books! FFS. But wait – it’s not only authors who may be interested in running an AI locally on their computers. AI can present tools for many use case scenarios. It’s all FREE. Right if that’s too good to be true, leave it alone. Go away and pay through the nose to the likes of ChatGPT. Well, I know the spiel, “I’m not an IT person. As for AI they just generate rubbish; never helped me.” If you’re one of those people, kindly depart and save yourself time for Instagram or Netflix, or wha’everrrr!

Whether or not you agree our Fat Disclaimer applies

In none of the following, I don’t need to know what each thing is. Most people who drive a car don’t need to understand all the electronics and parts of the engine that make it work. But may be you don’t drive a car, and won’t ever until you understand everything about how they work. That’s fine – you have individual choice! Like to leave now! This post is not meant to be helpful. It’s my notes. I put it here so I can find it in a flash if/when I need it in the future. If it helps one person, fine. If it helps nobody else, fine.

Running a local Large Language Model (LLM) for creative writing presents a unique challenge: balancing “model intelligence” against “memory capacity.” For an author, an AI that is brilliant but forgets the plot after three pages is useless.

This post details the technical process of optimizing a modern 12B parameter model for a system with 16GB RAM and a 10GB VRAM GPU. Do I know what all that means or understand it? I don’t. With AI assistance, the following was the work done. Yes – it takes time. What you want everything done instantly? The time spent will save me loads of time and headache as I expand my authorship career.

Caution: For idiots, AI software needs powerful computers and large storage. Their footprint can be 30 to 100GB. So not this stuff is not for ‘everybody’.

1. The Legacy Fallacy: Why Timestamps Matter

The initial candidate was a legacy quantisation of Mixtral 8x7B. While Mixtral was a landmark model, the specific files were tagged as “782 days old.” In the 2026 AI ecosystem, this represents a significant technical debt:

  • Architecture Gaps: Older GGUF files often lack support for modern features like Flash Attention or specialised KV cache quantisation.
  • Size vs. Resource Mismatch: The model weights (26.4 GB) far exceeded the available 10GB VRAM + 16GB RAM overhead, guaranteeing a “performance cliff” where the system would swap to disk, slowing generation to a crawl.

2. Model Selection: Finding the “Sweet Spot”

The goal was to find a model that maximises “literary reasoning” while fitting into a 10GB VRAM envelope.

  • Selected Model: Mistral-Nemo-12B-Instruct-v1
  • Quantisation: Q4_K_S (~7.12 GB)
  • Rationale: The 12B parameter count offers a significant upgrade in nuance over 7B or 8B models but is small enough to leave “room for memory” on a 10GB card. The Q4_K_S format provides a 4-bit quantisation that preserves high accuracy while keeping the base weight under 7.5 GB.

3. Hardware Triage: Reclaiming VRAM

A common pitfall is ignoring background VRAM usage. The test system initially showed 8.9 GB / 10 GB utilised before loading the model, likely due to open browsers or GPU-accelerated applications. Closing these reclaimed roughly 6 GB of VRAM, which was the difference between a successful load and an immediate crash.

4. The “Long-Term Memory” Tweak: KV Cache Quantisation

This is the most critical setting for authors. By default, the AI stores the “history” of the conversation (the KV Cache) at high precision (FP16).

  • The Tweak: Enabling 4-bit (Q4_0) KV Cache Quantisation.
  • Technical Result: This compresses the memory cost of every token by 75%.
  • Practical Impact: On a 10GB card, this allowed the Context Length to be safely increased from a measly 4,096 tokens (approx. 3,000 words) to 32,768 tokens (approx. 25,000 words). The AI can now “see” several chapters back into the story without exceeding the VRAM limit.

5. Software Configuration Summary

To replicate this author-centric environment in LM Studio (v0.4.0+), the following settings were applied:

SettingValuePurpose
GPU OffloadMax (40 Layers)Ensures the entire model runs on the fast graphics card.
Flash AttentionONReduces memory overhead during long-context processing.
Context Length32,768Provides enough memory for character and plot tracking.
Temperature0.8Encourages creative, varied prose over rigid assistant-like speech.
System PromptNovelist PersonaOverrides the default “Helpful Assistant” to focus on sensory details and “Show, Don’t Tell” methodology.

6. Monitoring the “VRAM Wall”

The finalised setup utilised 9.3 GB / 10.0 GB of VRAM. While this is “full,” the 4-bit KV cache ensures that generation speed remains high as the context window fills. If the usage hits 10.0 GB, the system is designed to spill the oldest tokens, maintaining stability for the user.

Technical Note to Self: Always check for newer model iterations (e.g., Mistral NeMo v2 or Llama 4 variants) as they become available, as the 700+ day gap in legacy models is usually indicative of obsolete quantisation logic.

Conclusion

Success! The model works fine. Now I can do lots of work offline and present a more distilled version for online AIs. That way I don’t burn through tokens online and get stopped out or invited to spend more money.

Others may find other tools for free online via LM Studio, for their own use case scenarios e.g. software development, graphic design etc.