

LLMs Explained Clearly: Training, Model Size, Quantization, and K-Quants
2026-05-18 · Manuel Spörer
From the outside, LLMs often feel like magic. You type in a question, get back a surprisingly useful answer, and quickly assume some kind of "understanding" must be at work. Technically, however, something far more sober is happening: a large language model calculates, step by step, which token best continues the existing context [1, 3].
That is precisely where modern LLMs draw their strength. They don't write convincingly because they "know" language the way a human does, but because they have learned statistical patterns from enormous amounts of text. Many of these models are built on the Transformer architecture, which processes relationships between tokens via self-attention and is therefore particularly well suited to language, long contexts, and complex dependencies [1, 2].
Anyone who wants to understand why local models suddenly appear in 4-bit variants, why a 32B model isn't automatically better than a smaller one, and what abbreviations like Q4_K_M actually mean has to cleanly separate a handful of fundamentals. That's what this article does: what LLMs are, how they are trained, which architectural differences matter in practice, and why quantization almost always plays a leading role for local AI.
What Are LLMs, Really?
Large Language Models, or LLMs, are large deep-learning models that can process, continue, and generate natural language because they have been trained on very large amounts of text [1]. Typically, an LLM doesn't produce a finished paragraph in one go but generates text sequentially: for the existing context, it calculates the most likely next token [1, 3].
A token is the fundamental computational unit of a language model. Depending on the tokenizer, it can be a whole word, a word fragment, a single character, or even a punctuation mark [3]. This seemingly minor technical detail matters because it explains why LLMs internally don't work with "words" or "meanings" in the human sense, but with numerical representations of token sequences.
The term "large" does not just refer to the file size of a model. It primarily means the number of learned parameters, the amount of training data, and the enormous compute effort needed to train such models in the first place [1, 4]. That's why LLMs can often handle tasks like summarization, translation, brainstorming, programming assistance, or drafting text surprisingly well: many of these problems can be framed as predicting the most likely token continuation [1].
But the sober view matters too. LLMs do not "know" what is true in any human sense. They generate plausible outputs based on learned patterns and the current prompt [1, 5]. That is precisely why they can deliver helpful answers — and a moment later produce convincingly worded errors, outdated information, or freely invented details [1, 5].
How and With What Are LLMs Trained?
The core of almost every modern LLM begins with pretraining. A model is trained on very large collections of text, code, and other documents that have previously been collected, cleaned, deduplicated, and preprocessed [1]. In classical causal language model training, the basic task is then surprisingly simple: predict the next token from a previous token sequence [3, 6].
This approach is often described as self-supervised learning because the target signal comes directly from the training material itself. The model doesn't need labels manually attached to each piece of text. The text supplies its own learning signal: the next token is the target [6].
Technically, training runs as a large optimization loop. The model makes a prediction, compares it to the actual next token, computes an error, and adjusts its weights via gradient methods so that the error shrinks across many training steps [6]. What looks like language competence from the outside is, at this level, first and foremost statistics, optimization, and a great deal of compute.
That is why the hardware question isn't a side topic. Large models are typically trained on many GPUs or other accelerators, because training and inference for large LLMs are extremely compute- and energy-intensive [1].
After pretraining, the work is usually not done. Fine-tuning often follows, adapting a base model to specific tasks, domains, response formats, or target behaviors [1]. An important special case is instruction tuning: the model learns to follow human instructions rather than just continuing raw text patterns [7].
A particularly well-known method in this context is RLHF — Reinforcement Learning from Human Feedback. For InstructGPT, OpenAI describes a multi-stage process consisting of supervised fine-tuning with human demonstrations, a reward model based on human preference comparisons, and a subsequent optimization step using reinforcement learning [7]. The goal is not just linguistic fluency but more helpful, more truthful, and less harmful responses [7].
Newer open models sometimes combine several of these stages: pretraining, supervised fine-tuning, reinforcement-learning phases, and additional model-based methods. OpenAI describes such a mix for gpt-oss, for example [8].
Are There Open Source LLMs — or Just Open Weight?
This is where a recurring confusion of terms kicks in. Many people say open source when they actually mean open weight. That isn't the same thing [9, 10].
According to the Open Source Initiative's definition, it isn't enough for software files to simply be downloadable. Open source covers, among other things, free redistribution, permitted modification, and non-discriminatory use [9]. For AI systems, the Open Source AI Definition goes further: it additionally requires transparent information about data, code, and the derivation of parameters so that a system is actually modifiable and verifiable [10].
That's precisely why many popular LLMs are, strictly speaking, open-weight models. Their weights are publicly available, but training data, data preparation, or the full training process often aren't fully disclosed [10].
Even so, there is now a growing set of openly available models that are highly relevant for practice, research, and local use. Examples include Qwen, MiniMax-M1, Mistral models, BLOOM, and OpenAI gpt-oss [8, 11, 12, 13, 23].
Alibaba Cloud describes Qwen3 as a model family with publicly available weights, including dense and Mixture-of-Experts variants ranging from 0.6B to 235B-A22B [11]. The associated repositories list Apache 2.0, and the Qwen3-235B-A22B model card also specifies apache-2.0 [11, 24].
MiniMax-M1 is described in the technical report as an open-weight model with a hybrid MoE architecture, lightning attention, 456 billion total parameters, and 45.9 billion active parameters per token [23]. According to MiniMax, M1 natively supports a 1-million-token context window, and the model weights are distributed through the official Hugging Face and GitHub accounts [25].
Mistral promotes its own open models for training, distillation, fine-tuning, and deployment [12]. BLOOM, in turn, is an autoregressive LLM from the BigScience project that, according to its model card, can generate text in 46 languages and 13 programming languages [13]. And with gpt-oss-120b and gpt-oss-20b, OpenAI has also released two open-weight language models under the Apache 2.0 license [8].
For practical purposes, a simpler question is often more useful than the labeling debate: what concretely is open here — just the weights, or also the data, code, license rights, and training documentation? That determines whether a model is merely usable, actually modifiable, or genuinely transparent and reproducible [8, 9, 10].
Which Architectural Differences Really Matter for LLMs?
When people talk about architecture, they don't mean marketing names but the internal makeup of a model: which Transformer-family components are used, how inputs are processed, and how the output text is produced [21].
A central distinction is between encoder-only, decoder-only, and encoder-decoder [21]. Encoder-only models are mainly designed for understanding and representing input text — for classification, search, embeddings, or text analysis [21]. Decoder-only models continue text autoregressively from left to right and are therefore the typical basis of modern chat and text generation models [21]. Encoder-decoder models combine both sides and are often used for sequence-to-sequence tasks like translation, summarization, or question answering [21].
Many of today's chat LLMs belong to the decoder-only family, because that architecture fits next-token prediction very well [3, 21]. This also explains why many capabilities of modern language models stem from the same underlying principle: read the context, compute the next token, repeat.
A second important distinction concerns dense models and Mixture of Experts, or MoE for short [8, 22]. In a dense model, essentially the same parts of the network are used for every prediction. In a sparse MoE model, a router decides per token which expert components become active [22].
Mistral describes Mixtral, for example, as a sparse Mixture-of-Experts network in which two out of eight expert groups are selected per token in each layer [22]. As a result, an MoE model can have a very large total parameter count while only activating a fraction of those parameters per token [8, 22]. OpenAI cites about 117 billion total parameters for gpt-oss-120b, but only 5.1 billion active parameters per token [8].
This is where an important practical point comes in: fewer active parameters do not automatically mean the model needs little memory overall. MoE primarily reduces the compute required per token. The expert weights generally still need to be available. An MoE model is therefore not automatically a RAM- or VRAM-saver. It can deliver a very strong ratio of model capacity to inference cost, while dense models are often more predictable, easier to operate, and easier to train.
In addition, LLMs differ via attention variants, maximum context length, the tokenizer, positional encoding, and various memory and inference optimizations [2, 16, 21]. Self-attention is at the heart of modern Transformers because it weights which tokens matter to which other tokens in the context [2]. The context length, in turn, determines how many tokens a model can consider at once — a decisive factor for long documents, chat histories, or large codebases [2, 8].
That's why two models with similar parameter counts can behave very differently in practice. Architecture, active parameters, context window, attention mechanics, and inference optimizations often shape the real-world profile of a model more than any single marketing number [8, 16, 21].
What Does Model Size or Parameter Count Mean?
When a model is labeled 4B, 32B, or 120B, the number usually refers to the approximate number of parameters — that is, 4 billion, 32 billion, or 120 billion learned numerical values inside the network [8, 11, 14]. These parameters determine how inputs are processed and how outputs are computed [14].
Bigger often means more capacity to represent patterns in the training data. But bigger does not automatically mean better [4, 7]. A model with more parameters can be more capable, but it can also be more inefficient, harder to operate, or simply oversized for a given task.
A well-known result from the InstructGPT work illustrates this: a 1.3B-parameter model with human feedback was preferred by human raters over a 175B-parameter GPT-3 model [15]. Model size is therefore just one dimension. Training objective, data quality, alignment, and fine-tuning can matter just as much [7, 15].
For practice, it also matters what size technically entails: higher memory consumption, higher hardware requirements, and often slower inference [8, 16].
For Mixture-of-Experts models, you need to look more carefully. OpenAI cites 117 billion total parameters for gpt-oss-120b but only 5.1 billion active parameters per token [8]. For gpt-oss-20b, the same source lists 21 billion total parameters and 3.6 billion active parameters per token [8]. Anyone comparing models should therefore look beyond the headline number and consider total parameters, active parameters, architecture, context length, quantization, and benchmarks [8, 16].
What Is Bit Quantization?
As soon as LLMs need to run locally, quantization comes up almost inevitably. The term refers to storing or processing a model's weights or activations at lower numerical precision [16, 17].
Instead of storing a model entirely in 32-bit or 16-bit floating point, weights can also be represented in lower formats like 8-bit or 4-bit [16, 17]. This reduces memory usage and can speed up inference because less data needs to be moved and processed [16, 17]. Hugging Face describes quantization in exactly this sense as a method to reduce memory and compute cost and to fit larger models into limited memory [16]. IBM describes it as the conversion of high-precision values like FP32 or FP16 into lower precision like INT8 [17].
The catch is the obvious one: less precision can cost quality [18]. Values are approximated more coarsely, and that approximation can affect answer quality, stability, or subtle aspects of model behavior. llama.cpp points out that quantization can shrink model weights and accelerate inference but can also introduce quality losses that can be observed, for example, via perplexity or Kullback-Leibler divergence [18].
A rough rule of thumb: the lower the bit count, the smaller the model — but the greater the risk that quality suffers [16, 18]. An 8-bit model usually stays closer to the original model but requires more memory [16, 17]. A well-engineered 4-bit model, by contrast, is often precisely the threshold at which local use on consumer hardware becomes practical at all [16, 18].
What Are K-Quants?
Anyone who downloads local GGUF models will sooner or later run into file names like Q4_K_M, Q5_K_M, or Q6_K. This is where the confusion starts for many.
K-Quants are quantization formats from the llama.cpp / GGML / GGUF ecosystem [18, 19]. Typical variants include Q2_K, Q3_K_M, Q4_K_S, Q4_K_M, Q5_K_M, and Q6_K [18, 19]. The Q essentially stands for quantization, the number roughly indicates the target precision in bits, and the K refers to the K-Quant family [19].
The suffixes S, M, and L usually stand for variants like small, medium, and large — that is, for different trade-offs between quality and memory within the same bit class [19]. The decisive point: K-Quants are not simply "everything rounded down to 4 bits." They use block and superblock structures along with scale values to balance memory consumption and quality loss more effectively [20].
In the llama.cpp discussion on QX_4, superblocks are described as groups of multiple quantization blocks that use quantized scales to keep the effective bits per weight low [20]. One example described there for 4-bit quantization works with superblocks of 16 blocks of 8 weights each, plus their own quantized scales, which effectively comes out to 5.125 bits per weight [20].
In practice this means: Q4_K_M is often a very useful sweet spot for local use. The model is heavily compressed but usually delivers better quality than older or simpler 4-bit quantizations [18, 20]. Q5_K_M needs more memory but can come closer to the quality of the unquantized model [18, 20]. Q6_K is larger still and is often chosen when quality matters more than minimum file size [18, 19].
For local LLMs, the choice between Q4_K_M, Q5_K_M, and Q6_K is therefore not a minor detail but a typical everyday trade-off between RAM or VRAM, speed, and answer quality [16, 18, 20].
Conclusion: It's Not Just Model Size — It's the Interplay
LLMs are large, mostly Transformer-based language models that are trained via next-token prediction and further shaped into useful response patterns through fine-tuning [1, 3, 7]. But their performance does not hinge on parameter count alone. Data quality, training procedure, architecture, alignment, inference stack, and quantization are at least as decisive [1, 7, 8].
For practice, that means: anyone working with local models should not only ask "how many billion parameters does this model have?" but also: which architecture does it use? How large is the context window? Is it dense or MoE? What quantization is it available in? And what license or degree of openness does it actually come with?
That's exactly where general AI hype turns into a workable technical assessment. Less magic, more mechanics. Less marketing, more system understanding. And that is the foundation for making sound judgments about local LLMs, hardware requirements, and quantization formats.
Sources
[1] IBM. What Are Large Language Models (LLMs)? https://www.ibm.com/think/topics/large-language-models
[2] Google Machine Learning Crash Course. LLMs: What's a large language model? https://developers.google.com/machine-learning/crash-course/llm/transformers
[3] Hugging Face Transformers Docs. Causal language modeling. https://huggingface.co/docs/transformers/en/tasks/language_modeling
[4] NVIDIA Developer Blog. An Introduction to Large Language Models: Prompt Engineering and P-Tuning. https://developer.nvidia.com/blog/an-introduction-to-large-language-models-prompt-engineering-and-p-tuning/
[5] OpenAI / Achiam et al. GPT-4 Technical Report. https://arxiv.org/abs/2303.08774
[6] Hugging Face LLM Course. Training a causal language model from scratch. https://huggingface.co/learn/llm-course/chapter7/6
[7] OpenAI. Aligning language models to follow instructions. https://openai.com/index/instruction-following/
[8] OpenAI. Introducing gpt-oss. https://openai.com/index/introducing-gpt-oss/
[9] Open Source Initiative. The Open Source Definition. https://opensource.org/osd
[10] Open Source Initiative. The Open Source AI Definition 1.0. https://opensource.org/ai/open-source-ai-definition
[11] QwenLM / GitHub. Qwen3. https://github.com/QwenLM/Qwen3
[12] Mistral AI. Frontier AI LLMs, assistants, agents, services. https://mistral.ai/
[13] BigScience / Hugging Face. BLOOM Model Card. https://huggingface.co/bigscience/bloom
[14] Google Machine Learning Crash Course. Introduction to Large Language Models. https://developers.google.com/machine-learning/crash-course/llm
[15] Ouyang et al. Training language models to follow instructions with human feedback. https://arxiv.org/abs/2203.02155
[16] Hugging Face Transformers Docs. Quantization. https://huggingface.co/docs/transformers/en/main_classes/quantization
[17] IBM. What is Quantization? https://www.ibm.com/think/topics/quantization
[18] llama.cpp. Quantize README. https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md
[19] llama.cpp quantization table mirror with K-Quant BPW examples. https://gitlab.informatik.uni-halle.de/ambcj/llama.cpp/-/blob/3e916a07ac093045d88ef0c4fa78647ae0efc010/examples/quantize/README.md
[20] ggml-org / llama.cpp GitHub Issue. QX_4 quantization. https://github.com/ggml-org/llama.cpp/issues/1240
[21] Hugging Face LLM Course. Transformer Architectures. https://huggingface.co/learn/llm-course/chapter1/6
[22] Mistral AI. Mixtral of experts. https://mistral.ai/news/mixtral-of-experts
[23] MiniMax et al. MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention. https://arxiv.org/abs/2506.13585
[24] Qwen / Hugging Face. Qwen3-235B-A22B Model Card. https://huggingface.co/Qwen/Qwen3-235B-A22B
[25] MiniMax. MiniMax-M1, the World's First Open-Source, Large-Scale, Hybrid-Attention Reasoning Model. https://www.minimax.io/news/minimaxm1