Transformers Runs GGUF Quants With llama.cpp Kernels

Transformers Now Runs llama.cpp GGUF Quants on Apple Silicon

By Indie Kings | October 10, 2026

Updated October 10, 2026: Transformers can now load and run GGUF quants directly in PyTorch through ggml kernels supplied by the kernels library. This article explains the Apple Silicon first release around Qwen3.5, the install and load path, quant sizes, serve options, benchmarks, limits, and practical use cases, per the Hugging Face blog article Transformers now runs llama.cpp quants by Marc Sun, Arthur Zucker and Lysandre Debut dated Sep 22 2026.

Standing: Per HF blog reporting, no separate store listing applies.

Transformers GGUF quants thumbnail

Image: Transformers GGUF quants thumbnail. Credit: Hugging Face blog.

What is new

Per the Hugging Face blog article Transformers now runs llama.cpp quants, Transformers can now run GGUF quantized checkpoints without leaving PyTorch. The work routes GGUF weights through ggml kernels that ship through the kernels library. The initial focus is Apple Silicon and Qwen3.5. That focus matters because Apple Silicon laptops are common local inference machines, and Qwen3.5 gives both dense and Mixture of Experts coverage in one first target.

The same Hugging Face blog post frames this as a bridge, not a replacement. The llama.cpp project remains the recommended engine for efficient local inference. The new Transformers path is for people who want the GGUF weight ecosystem inside a PyTorch workflow. The distinction is important for expectations. If raw speed and the smallest memory footprint are the only goals, the Hugging Face authors still point to llama.cpp. If experiment control, evaluation, validation, custom decoding, or fine tuning preparation are the goals, the new path adds value.

The announcement names Marc Sun, Arthur Zucker and Lysandre Debut as authors and carries a Sep 22 2026 date. This Indie Kings article uses only facts from that Hugging Face blog article and attributes them explicitly throughout. Every number, filename style option, version string, model name, and behavior below comes Per HF reporting in that post. Where this article restates an implication, it stays inside the scope of the post and does not add outside measurements or outside model claims.

The core idea is simple to state and useful to keep in mind while reading the rest of this page. GGUF files hold quantized weights in a compact form. Transformers historically wanted full precision or its own quantized formats in Torch. Now Transformers can accept a GGUF file reference at load time, bring up matching ggml kernels through the kernels library, and run generation. On the supported Metal packed path, the stack also brings in ggml attention support automatically. Outside that path, it falls back to scaled dot product attention with a warning, Per HF.

Apple Silicon is the launch platform, not the final boundary. The Hugging Face post describes a Metal packed path that loads ggml and Metal kernels plus ggml attention. That sentence tells a local Mac owner what to expect. On a supported Mac, the fast path should engage without manual kernel selection. On other setups, the same GGUF checkpoint can still load, but attention falls back, with a warning, to a more general implementation. Readers should treat that fallback as a compatibility safety net, not as the performance target.

Qwen3.5 is the launch model family, with dense and Mixture of Experts variants covered, plus compatible Qwen3.8 support noted as the only extra. That scope keeps the first release testable. It also sets the vocabulary for the rest of this article. When the post says start with Q4_K_M, it means the Qwen3.5 4B quant menu discussed below. When it lists future work around padding and batching, it means current single stream local runs are the sweet spot and batched server style loads are still pending.

How it works

Per HF, the implementation runs GGUF through ggml kernels exposed through the kernels library. In practical terms, the GGUF file stays in its quantized layout, and the compute kernels understand that layout. There is no requirement in the post to convert the checkpoint to another format before inference. That is the point of the feature. A researcher can point Transformers at a GGUF quant, keep the file size benefit of quantization, and still use familiar Transformers generation calls.

The Metal packed path is the optimized route on Apple Silicon. Per HF, it auto loads ggml and Metal kernels and ggml attention. Auto loading matters because kernel wiring is often the hardest part of a new backend. The user does not pick a kernel per layer in the described flow. The loader detects the supported path and brings the right pieces. That design reduces setup friction for Mac owners who simply want to try a Q4_K_M checkpoint and compare it with full precision behavior.

The fallback path is also explicit. When the Metal packed path is not available, the stack falls back to scaled dot product attention, described in the post as sdpa, with a warning. A fallback with a warning is honest engineering. It preserves the ability to run, inspect outputs, test prompts, and validate logic on machines that lack the optimized route. It also signals that speed on the fallback path should not be compared directly with the packed path or with llama.cpp decode only numbers.

The kernels library is the delivery vehicle. The Hugging Face post says to install the main branch of Transformers plus kernels. That pairing tells readers this was fresh work at publication time, newer than the last stable Transformers release tag. It also explains the benchmark version pins later in this article. The post pins kernels to version 0.17.0 alongside PyTorch 2.12.1 for its MacBook Pro measurements. Those pins help future readers interpret whether a newer install should behave the same.

Attention handling deserves a second paragraph because attention often decides local speed. On the packed path, ggml attention is part of the auto loaded set. Off that path, scaled dot product attention covers correctness. This split is common in early hardware specific releases. Optimize the hot path first, keep a portable path for everything else, warn when the portable path engages so benchmarks are not misread. Readers who test on Apple Silicon should confirm which path engaged before drawing conclusions about tokens per second.

The headline takeaway for this section is architectural. Transformers remains the orchestration layer for tokenization, model structure, caching, stopping logic, and serving. The ggml kernels supply the math for quantized weights. The kernels library supplies packaging and dispatch. None of this changes the Transformers programming model described in the post. Loading names a GGUF file, generation uses generate, serving uses the Transformers serve command, and fine tune preparation uses a dequantize option through GgufConfig.

Quant sizes

Per HF, the Qwen3.5 4B family spans a wide size range from full precision to small quants. Full precision BF16 is listed at 8.42GB. The Q6_K quant is listed at 3.53GB. The Q5_K_M quant is listed at 3.14GB. The Q4_K_M quant is listed at 2.74GB. Those four numbers tell the whole storage story in one glance. Moving from BF16 to Q4_K_M cuts the file to roughly one third of the full precision size.

The Hugging Face post advises starting with Q4_K_M. That advice is practical for local laptops. Q4_K_M is the smallest listed option in the table, and therefore the easiest fit for limited unified memory. The post also notes that Q4_K_M mixes precisions. Mixed precision inside a K quant means different tensors or blocks can use different bit widths to protect quality where it matters most while saving space elsewhere. Readers should not read Q4_K_M as a single uniform bit width across the whole model.

Size is not the only variable, but it is the first filter for local runs. A Mac with 32GB of unified memory, like the benchmark machine in the post, can hold larger options comfortably. A Mac with less memory may still favor Q4_K_M for headroom. The post does not promise identical quality across quants. It positions the feature as a way to experiment in PyTorch, evaluate quants, and validate conversions. In other words, load two sizes, run the same prompts, compare outputs, and decide which tradeoff fits a given task.

Community publishers such as Unsloth also distribute quantized checkpoints in the broader ecosystem, and Apple Silicon owners often collect several quant levels before choosing one. This new Transformers path fits that habit. Keep the small quant for quick iteration, keep a larger quant for quality checks, and use the same Transformers code for both. The Hugging Face framing around evaluation and validation supports exactly that side by side style.

The table below repeats the four Per HF numbers for Qwen3.5 4B so readers can cite them cleanly. All sizes are in gigabytes as reported in the Hugging Face blog post. The order runs from largest to smallest to make the savings obvious. Use this table when planning downloads, disk space, and memory headroom.

Checkpoint formatReported sizeWhat it means for local runs
BF168.42GBFull precision reference, largest footprint
Q6_K3.53GBQuantized option, less than half of BF16
Q5_K_M3.14GBSmaller quantized option for tighter memory
Q4_K_M2.74GBSmallest listed option, recommended starting point, mixes precisions

The second paragraph after the table matters for interpretation. These sizes describe the Qwen3.5 4B quant menu in the Hugging Face post. They do not describe every model or every context size. Larger parameter counts and longer contexts change memory needs. Treat the table as a concrete example of GGUF savings, not as a universal claim about all quants. The broader point stands across models though. GGUF quants exist to make strong models fit local machines, and Transformers can now consume them directly.

Install and load

Per HF, installation uses the main branch of Transformers plus the kernels library. The post frames this as a pip style install from main rather than a stable version install. Readers in October 2026 should check whether the feature has since reached a stable release, but at the Sep 22 2026 publication date, main was the stated route. The kernels library version pinned for benchmarks is 0.17.0, which gives a useful reference point when reproducing results.

Loading names a GGUF file through the pretrained loader. Per HF, the call uses from pretrained style loading with a gguf file argument. That design keeps the rest of the Transformers workflow unchanged. Tokenizers, chat templates, generation settings, and stopping logic stay where users expect them. Only the weight source changes from a full precision checkpoint to a GGUF quant file. For people who already script Transformers evaluations, this minimizes code churn.

On Apple Silicon, the Metal packed path engages automatically when available. Per HF, that path loads ggml and Metal kernels and ggml attention. Automatic engagement is good for first time users. It removes a class of configuration mistakes where the wrong kernel gets selected for the wrong device. It also means benchmark readers should record whether the packed path engaged, because the fallback path uses scaled dot product attention with a warning and will not represent peak speed.

Dequantization is available for training preparation. Per HF, fine tuning preparation goes through GgufConfig with a dequantize option. That option matters because quantized weights are efficient for inference but are often expanded before gradient updates. The post lists fine tuning through dequantization as one of the core use cases, alongside experimentation, evaluation, validation, and custom decoding. In short, load the quant to test ideas, then dequantize when it is time to train.

A careful loader also helps validation workflows. Suppose a team converts a checkpoint to GGUF and wants to confirm the conversion is sane. Loading that GGUF file back into Transformers and running fixed prompts gives a direct check. Tokenizer behavior, template behavior, and decoding behavior stay inside one framework, so differences can be attributed to the quant rather than to a change of engine. The Hugging Face post explicitly names validating conversions as a use case.

The install and load story is therefore short on purpose. Install main plus kernels, point the pretrained loader at a GGUF file, let the packed path load the right kernels on Apple Silicon, watch for the scaled dot product attention warning elsewhere, and use dequantization through GgufConfig when moving from inference to fine tuning. That sequence covers the whole getting started flow Per HF without extra tooling.

Serve API

Per HF, Transformers serve exposes an OpenAI compatible API on localhost port 8000. Compatibility here means existing HTTP clients built for chat completions style endpoints can point at a local server with minimal changes. For local developers, that is often more convenient than writing a custom loop. Start the server, send prompts over HTTP, receive completions, and keep the model resident in memory across requests.

The serve command includes a reasoning control with auto, on, and off modes. Per HF, the option is named as a reasoning flag and accepts those three values. Auto lets the server decide based on model capability or request context. On forces reasoning style output handling where supported. Off disables it. Readers should describe the control in words rather than copying shell dashes into prose, because the exact command spelling can evolve after the Sep 22 2026 post. The behavior concept is stable though. Reasoning output needs explicit control in an API.

Localhost port 8000 is a familiar default for development servers. It keeps traffic on the machine, avoids exposing a test endpoint to a network, and matches many example clients. The Hugging Face framing presents serving as part of the same feature, not as a separate product. Load a GGUF quant, serve it locally, and exercise it through a standard API shape. That loop is useful for app prototyping, prompt iteration, and regression checks across quants.

Apple Silicon owners benefit most at launch because the packed path and the server path compose. The fast kernels handle the math, while the server handles batching of requests at the HTTP layer as implemented. Current padding and batching work is still listed as pending in the post, so readers should keep expectations modest for heavy concurrent loads. Single user iteration, demos, and evaluation harnesses are the natural fit for this first release.

The serve path also helps teams compare quants fairly. Serve Q4_K_M, run a fixed prompt set, record outputs and latency impressions, then repeat with Q5_K_M or Q6_K. Because the API shape stays constant, the client code does not change between runs. Only the model file and the server restart change. That isolation is valuable when deciding whether a smaller quant preserves the behaviors a product needs.

Benchmarks

Per HF, benchmarks run on a MacBook Pro with M2 Max and 32GB of memory, running macOS 26.6, PyTorch 2.12.1, and kernels 0.17.0. Those pins are worth repeating because local inference numbers move with operating system, Torch version, and kernel version. A reader who tests with a different stack should expect some drift. The post is careful to name the full stack, which makes its chart easier to interpret honestly.

The headline result is that Transformers performance is close to llama.cpp on the tested setup. That is a meaningful claim for a first release. llama.cpp is a mature, highly tuned runtime for GGUF. A new PyTorch side path that lands near it suggests the ggml kernel integration is sound. The Hugging Face authors still recommend llama.cpp as the efficient local inference engine, so closeness should be read as validation, not as a claim of superiority.

The comparison caveat is explicit and important. Per HF, the chart does not use identical conditions. The Transformers numbers include prefill, while the llama.cpp numbers are decode only. Prefill covers processing the prompt, while decode covers generating new tokens. Including prefill adds work that decode only numbers omit, so the two bars measure overlapping but not identical phases. Readers should not treat the chart as a head to head race. They should treat it as evidence that the new path is in a sensible range.

Local benchmark hygiene still applies. Close background apps, use the same prompt lengths across runs, record whether the Metal packed path engaged, note whether any scaled dot product attention warning appeared, and keep prompts and generation settings fixed when comparing Q4_K_M against larger quants. The Hugging Face post does not ask readers to take one chart as final. It gives the setup pins so others can rerun and refine.

Memory behavior also matters as much as tokens per second. The quant size table shows why. A 2.74GB Q4_K_M file leaves far more headroom than an 8.42GB BF16 file on the same 32GB machine. Headroom affects context length, cache comfort, and multitasking. The post does not turn that observation into a hard performance promise. It simply gives the sizes and the speed context together so Mac owners can plan.

The benchmark section therefore ends where it began. Same machine, same stack pins, close to llama.cpp, with a stated measurement mismatch around prefill versus decode only. That is a careful first benchmark. It informs without overclaiming, and it invites reproduction. Future posts may tighten the methodology, but this release establishes that GGUF in Transformers on Apple Silicon is already practical.

Use cases

Per HF, the post names five practical use cases. Experiment in PyTorch, evaluate quants, validate conversions, build custom decoding, and prepare fine tuning through GgufConfig dequantization. Those five cover the full research loop from first load to training. They also explain why someone would choose Transformers over llama.cpp for a given task even though llama.cpp remains the efficiency recommendation.

Experimenting in PyTorch means staying inside a familiar research stack. Data loading, metric code, prompt templates, and analysis notebooks do not need to move. A GGUF quant becomes one more checkpoint format that the same scripts can open. That continuity lowers the cost of trying a smaller quant. It also makes results easier to share, because reviewers and teammates already understand the Transformers side.

Evaluating quants means running the same tasks across Q4_K_M, Q5_K_M, Q6_K, and BF16 and recording where quality drops. The size table makes the tradeoff concrete. Each step up in size costs disk and memory but may recover accuracy, instruction following, reasoning stability, or long context coherence. The Hugging Face framing does not pick a winner. It gives the loading mechanism and advises starting at Q4_K_M, then leaves the judgment to task specific tests.

Validating conversions means checking that a newly produced GGUF file behaves as expected. Conversion bugs can be subtle. A tokenizer mismatch, a template mismatch, or a tensor mapping slip can look like a quality problem when it is really a packaging problem. Loading the GGUF back into Transformers isolates the package from the engine. If the same prompts behave sensibly in Transformers, the file is likely sound. If they do not, the conversion needs another look.

Custom decoding is the fourth use case and often the most creative. Researchers may want constrained outputs, speculative helpers, custom stopping rules, tool call shaping, or logging at each step. Transformers exposes those hooks in Python. Running GGUF weights under those hooks combines the efficiency of quants with the flexibility of research code. The generate loop fixes in the next section make this use case stronger, because stopping and masking behavior directly affect custom loops.

Fine tuning preparation through dequantization closes the loop. Per HF, GgufConfig supports a dequantize step for fine tuning. The typical pattern is to explore cheaply with quants, select a promising direction, expand weights as needed, and then train. Naming Unsloth style community quant publishers alongside Hugging Face hosted files helps here too. Whatever the source of a GGUF file, the ability to inspect it in Transformers before committing training budget reduces risk.

Beyond GGUF

Per HF, the ggml kernels matter beyond GGUF files. The same kernel family can support models or modalities that lack full Transformers support today. That sentence points to a broader strategy. GGUF is the first visible application, but the kernels library is the durable piece. Once low level math for quantization, normalization, attention, and ranking exists in a reusable form, new model types become easier to enable.

The post includes a kernels table with five entries. The names are ggml quantization, ggml norm, ggml attention, ggml gated delta net, and topk. Each name maps to a stage of model execution in broad terms. Quantization handles compact weights. Normalization stabilizes activations. Attention routes context. Gated delta structures support newer sequence mixing designs. Topk supports selection and routing steps used in sparse and retrieval style logic. The post does not expand each kernel into a tutorial, so this article keeps the descriptions at that architectural level.

The table below lists the five kernel names Per HF so readers can reference them precisely. Spelling matters for search and for code lookup. The short gloss in the right column stays generic on purpose, because the Hugging Face post scopes the release to Qwen3.5 and compatible Qwen3.8 and does not promise that every kernel is exercised by every model.

Kernel entryRole in the stack
ggml quantizationMath for quantized weight formats including GGUF quants
ggml normNormalization support for stable activations
ggml attnAttention support, auto loaded on the Metal packed path
ggml gated delta netSequence mixing support for newer model designs
topkSelection support for ranking and sparse routing steps

The strategic read is straightforward. A kernels library that already covers quantization, norm, attention, gated sequence mixing, and topk selection is positioned to help with unsupported models and modalities over time. The Hugging Face post states that direction explicitly without overpromising dates. Readers should treat the current Qwen3.5 plus compatible Qwen3.8 scope as what works now, and the kernel list as a hint about where enablement may go next.

Apple platform work and community quant work both benefit from shared kernels. Apple Silicon owners get the packed Metal route first. Publishers in the style of Unsloth and teams publishing to Hugging Face hubs supply the files. Transformers plus kernels sits in the middle as the common runtime for checks. That triangle of hardware, files, and runtime is why a kernel table belongs in a GGUF announcement at all.

Generate loop fixes

Per HF, the release also fixes two generate loop issues. The first fix drops an unnecessary attention mask, referenced as pull request 48814. The second fix defers the stopping check, referenced as pull request 47975. Both sound small, but generate loops run once per token, so small per step savings and correctness tweaks compound across long outputs.

Dropping an unnecessary attention mask reduces redundant work. Masks tell attention which positions may be read. When a mask is provably unnecessary for a given path, skipping its construction or transfer saves time and memory traffic. The Hugging Face post does not quantify the saving in isolation. It lists the fix as part of making generation cleaner for the GGUF path and for Transformers generation more generally.

Deferring the stopping check changes when the loop tests for completion. A stopping check decides whether generation should halt because an end token appeared or a limit was reached. Checking too eagerly can add overhead or complicate custom stopping logic. Deferring it to a better point keeps the hot loop tight. For custom decoding researchers, this kind of change matters because their own stopping hooks interact with the built in check.

The pull request numbers give readers a trail to follow. Reference 48814 identifies the attention mask change. Reference 47975 identifies the stopping check change. Those numbers are Per HF and are included here for traceability, not as outside claims. Readers who want the diff level detail should look up those references against the Transformers repository history around the Sep 22 2026 post.

Taken together, the two fixes reinforce the use cases above. Faster, cleaner generation helps evaluation harnesses that run many prompts. More predictable stopping helps custom decoding and server side behavior. Neither fix changes the quant size math, but both improve the experience of running those quants inside Transformers rather than in a separate engine.

Limits

Per HF, the Metal packed path is listed as MPS only. MPS is the Apple graphics and compute route on macOS, so this limit restates the Apple Silicon first focus in platform terms. Mac owners are inside the optimized envelope. Other GPUs and operating systems can still use the fallback, with the scaled dot product attention warning, but they should not expect the packed path benefits described for Mac.

Padding and batching work is listed as pending. That phrase signals that single stream correctness came first and efficient multi sequence handling is still in progress. Local chat, single prompt evaluation, and small scale tests fit well now. High concurrency serving and large batched evaluation are better deferred until that follow up lands. Readers planning throughput tests should keep batches small and document settings.

Model scope is narrow at launch. Per HF, the supported set is Qwen3.5 dense and Mixture of Experts variants plus compatible Qwen3.8 only. That is a tight allow list. It keeps quality control manageable and avoids implying that every GGUF file on the internet will load. Owners of other model families should wait for expanded coverage rather than assuming the loader will accept arbitrary files.

The efficiency guidance is also a limit in spirit. The Hugging Face authors state that llama.cpp remains the recommended engine for efficient local inference. That sentence bounds the whole release. Transformers with ggml kernels is the flexible research path. llama.cpp is the lean deployment path. Choosing between them is a tradeoff between control and peak efficiency, not a contest with one winner.

These limits read as healthy scoping rather than as warnings. An MPS only packed path, pending batching work, and a short model allow list are normal for a first integration of a new weight format. The post discloses them plainly, pins benchmark versions, and flags the chart mismatch between prefill inclusive and decode only numbers. That transparency makes the feature easier to trust and easier to retest as coverage grows.

FAQ

What did Hugging Face announce about Transformers and GGUF quants?

Per the Hugging Face blog article Transformers now runs llama.cpp quants by Marc Sun, Arthur Zucker and Lysandre Debut dated Sep 22 2026, Transformers can run GGUF quants through ggml kernels supplied by the kernels library, with an initial focus on Apple Silicon and Qwen3.5.

Which Qwen3.5 4B sizes are reported and where should I start?

Per HF, BF16 is 8.42GB, Q6_K is 3.53GB, Q5_K_M is 3.14GB, and Q4_K_M is 2.74GB. The post advises starting with Q4_K_M, which mixes precisions to balance size and quality.

How do I install and load a GGUF quant Per HF?

Install the main branch of Transformers plus the kernels library, then use pretrained style loading with a GGUF file argument. On Apple Silicon the Metal packed path auto loads ggml and Metal kernels plus ggml attention, and elsewhere it falls back to scaled dot product attention with a warning.

How does local serving work?

Per HF, the Transformers serve command exposes an OpenAI compatible API on localhost port 8000. A reasoning control supports auto, on, and off modes for handling reasoning style output.

What benchmark setup and caveat should I cite?

Per HF, tests used a MacBook Pro with M2 Max and 32GB of memory on macOS 26.6 with PyTorch 2.12.1 and kernels 0.17.0, with results close to llama.cpp. The chart is not apples to apples because Transformers numbers include prefill while llama.cpp numbers are decode only.

What are the limits and is llama.cpp still recommended?

Per HF, the packed path is MPS only, padding and batching work is pending, and support covers Qwen3.5 dense and Mixture of Experts plus compatible Qwen3.8 only. Yes, llama.cpp remains the recommended engine for efficient local inference, while the Transformers path serves experimentation, evaluation, validation, custom decoding, and fine tuning preparation through GgufConfig dequantization.

Bottom Line

Transformers can now run GGUF quants in PyTorch through ggml kernels from the kernels library, starting with Apple Silicon and Qwen3.5. Start with the 2.74GB Q4_K_M quant, use the pretrained loader with a GGUF file argument, expect the Metal packed path to auto load ggml attention on Mac, and accept the scaled dot product attention fallback with a warning elsewhere. Use local serving on localhost port 8000 with the reasoning control set to auto, on, or off as needed. Cite the MacBook Pro M2 Max benchmark pins of macOS 26.6, PyTorch 2.12.1, and kernels 0.17.0, remember the prefill inclusive versus decode only caveat, and keep llama.cpp as the efficiency pick while using Transformers for experiments, quant evaluation, conversion validation, custom decoding, and GgufConfig dequantization for fine tuning. Per HF reporting, the packed path is MPS only, batching work is pending, and model scope is Qwen3.5 dense and Mixture of Experts plus compatible Qwen3.8 only.

Sources

Per HF, all technical facts in this article come from the Hugging Face blog article Transformers now runs llama.cpp quants by Marc Sun, Arthur Zucker and Lysandre Debut dated Sep 22 2026. The named sources in prose include Hugging Face, Apple, and Unsloth without added hyperlinks, and the only external link on this page is the source URL: Transformers now runs llama.cpp quants.

Share