Skip to content
#

int4

Here are 77 public repositories matching this topic...

Qwen3.8-Flash-Next on 2× RTX 3090: with 128 GB RAM up to 4,191 tok/s prefill · 111.5 tok/s decode (131K prompt), full 256K window at 2,865 tok/s; with 64 GB RAM 3,410 tok/s prefill · 84 tok/s decode. New: opt-in uncensored mode (runtime abliteration, no new weights).

  • Updated Oct 2, 2026
  • Python

600 KB WASM runtime for Cactus Compute's Needle AI tool-calling models, Needle 3, 2 and 1 from one build. Browser, Cloudflare Workers, Node.js, Python, C FFI and no_std. Token-exact with the JAX reference. No backend, no API key.

  • Updated Sep 19, 2026
  • Rust

Serving 4-bit Qwen3.8-27B on a single DGX Spark (GB10): 75 tok/s single-stream, 246 tok/s aggregate at 8-way concurrency. NVFP4 vs MixedInt4-AutoRound vs the FP8 baseline, measured on one harness — including why the quantization advantage collapses to +0.2% by c16.

  • Updated Sep 29, 2026
  • Python

⚡️ The fastest way to run local LLMs on Apple Silicon — sub-second model loads, beats Ollama on throughput, tail latency, and full-response time. OpenAI/Ollama-compatible. No cloud, no API keys.

  • Updated Oct 3, 2026
  • Python

Research and training stack for AVA — a tool-using, memory-aware virtual assistant targeting 4 GB VRAM. Spans custom transformers, verifier-RL, external memory, multi-domain benchmarks, and Gemma 4 inference optimization.

  • Updated Sep 8, 2026
  • Python

Run Qwen3.8-Flash-Next on ONE RTX 3090 (24 GB) + 64 GB RAM: 128K context, up to 2,100 tok/s prefill, 43–51 tok/s decode. vLLM runtime with hot MoE experts on the GPU and cold experts computed on the CPU, INT8 KV cache, Docker, OpenAI-compatible API.

  • Updated Sep 30, 2026
  • Python

Add this topic to your repo

To associate your repository with the int4 topic, visit your repo's landing page and select "manage topics."

Learn more