Skip to content
#

llm-benchmark

Here are 384 public repositories matching this topic...

[ICML 2024] TrustLLM: LLM trustworthiness evaluation across truthfulness, safety, fairness, robustness, privacy and ethics. Python/CLI toolkit for local models, compatible APIs and AI agent orchestration.

  • Updated Oct 4, 2026
  • Python
awesome-ai-tokenomics

A curated list on AI token economics: what tokens cost, where they get wasted, and how to cut the bill. Tools, benchmarks, papers, and copy-paste configs for the token economy of LLMs and coding agents.

  • Updated Oct 4, 2026
  • Python
gcf

The AI-native wire format for structured data. 100% comprehension on every frontier model. 50-92% fewer tokens than JSON. 43B+ lossless round-trips across 17 formats. Spec v3.5.1 Stable.

  • Updated Oct 7, 2026
prism

Открытый бенчмарк LLM: какая нейросеть лучше пишет код 1С:Предприятие (BSL). Объективная оценка LLM по методике SMOP с реальным исполнением в 1С — Claude, GPT, Gemini, DeepSeek, YandexGPT, GigaChat.

  • Updated Oct 5, 2026
  • Python

Measured llama.cpp benchmarks on AMD Radeon RDNA4 with ROCm: RX 9070 XT + Radeon AI PRO R9700 (48 GB). Qwen3.8, Gemma 4, GLM, gpt-oss; GGUF quant quality (KLD), MTP speculative decoding, 128K context, Claude Code with local models.

  • Updated Oct 7, 2026
  • Shell

Add this topic to your repo

To associate your repository with the llm-benchmark topic, visit your repo's landing page and select "manage topics."

Learn more