Comparing changes

* Fix model download in test workflow * Use hf CLI in test workflow * Use hf CLI name in CI and docs * Reference PR in changelog

abetlen#2150) * fix(ci): use supported macos runner label * fix(ci): add apple silicon macos test coverage * fix(ci): run standard macos tests on apple silicon * fix(ci): simplify apple silicon macos install * fix(ci): disable ggml native on apple silicon runner * docs: update changelog for macos ci runner fix

* Add Ruff formatting and safe lint baseline * Update changelog for Ruff setup

* Update llama.cpp and sync bindings * Clean up binding compatibility shims * Remove flash attention property shim * Remove mtmd verbosity shim * Add docstrings for new bindings * Format Ruff files and add changelog entry

* ci: add riscv64 wheel builds to release workflow Add a build_wheels_riscv64 job mirroring the existing arm64 QEMU-based build. Uses cibuildwheel with QEMU emulation for linux/riscv64, targeting CPython 3.10-3.14 on manylinux. Closes abetlen#2138 * ci: use cibuildwheel 3.1.2 for riscv64 wheels * docs: update changelog for riscv64 wheel PR --------- Co-authored-by: abetlen <abetlen@gmail.com>

* fix: handle Qwen 3.5 hybrid prefix reuse * test: fix Qwen runtime unit mocks * test: drop Qwen runtime unit tests * docs: credit Qwen fix contributors in changelog * docs/tests: update default Qwen model to 3.5 0.8B * test: rebaseline Qwen 3.5 outputs * test: stabilize low-level Qwen sampling check * test: tighten Qwen 3.5 completion prompts

* fix(ci): harden release wheel workflow * fix(ci): document and pin release wheel baselines * fix(ci): speed up release arch builds * fix(ci): split riscv64 by python version * fix(ci): sanitize riscv64 artifact names

* fix(ci): harden cuda wheel workflow * fix(ci): pin cuda toolkit versions accurately * fix(ci): resolve exact cuda toolkit installs * fix(ci): align cuda toolkit roots and tags * fix(ci): pin cuda packages to nvidia label * fix(ci): allow cuda solver to mix non-cuda deps

* fix(ci): harden docker build workflow * docs: update changelog for ci workflows

* feat: expose attention_type parameter in Llama.__init__ * docs: preserve attention_type in pickled state * docs: update changelog for attention_type --------- Co-authored-by: Victor Biederbeck <victor@moria.hiddencove.xyz> Co-authored-by: abetlen <abetlen@gmail.com>

…ent arches and one PTX target for forward compatibility (abetlen#2158) * fix(ci): shrink CUDA wheel fatbins * docs: update changelog for cuda wheel size fix

* Fix embedding models without KV memory * Add changelog entry for embedding memory fix

* Update llama.cpp to c0159f9c1 * Add changelog entry for llama.cpp update

…#2165) * fix(ci): publish distinct manylinux and musllinux cpu wheels * docs: add changelog entry for linux wheel repair fix

* ci: publish CPU wheels as py3-none * docs: add changelog entry for py3-none wheel tags

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Comparing changes

Open a pull request

Uh oh!

Commits on Mar 22, 2026

Commits on Mar 23, 2026

Commits on Mar 24, 2026

Commits on Mar 25, 2026

Commits on Mar 29, 2026

Commits on Mar 30, 2026

This comparison is taking too long to generate.

Uh oh!