← 返回发现
local-benchmark-runner-public
AI工具ai

local-benchmark-runner-public

raydestar/local-benchmark-runner-public

Contamination-resistant LLM evaluation harness: sealed benchmark banks with fabricated evidence, hashes published before measurement, verified task completion, and family-level statistics. Works with any OpenAI-compatible endpoint.

0
Stars
15
热度评分
+0
7日增长
1天
趋势

📋 项目信息

分类AI工具
用途ai
发现日期2026-07-29