Concurrency gateway converting web-based AIs such as LMArena and Gemini into OpenAI-compatible APIs via Camoufox anti-detection browsing with multi-window concurrency and account isolation.
Microsoft cross-platform sandboxed execution container provides policy-driven layered isolation for untrusted code, offering multiple backends and containment controls for security-conscious developers and platforms.
Local kanban web app that creates isolated worktrees for each task to run CLI agents in parallel, managing dependency chains and supporting diff review.
Supplies a CLI and framework for creating, testing, and measuring Agent Skills, enabling skill authors to benchmark quality and compare effectiveness across different models.
Desktop application for browsing, installing, and managing AI agent skills, supporting more than twenty agents with local-first storage and skill synchronization.
High-performance CUDA kernel library implementing Kimi Delta Attention to accelerate linear-attention computation during large-language-model inference for researchers and engineers optimizing throughput.
The official implementation of Agent-as-a-Router for coding tasks, providing benchmark data, reference outputs, and gateway-level integration demos for backend model routing.
A library providing environments and infrastructure for evaluating and improving models and agents at scale. Supports reproducible benchmarks and reinforcement-training integrations across multiple frameworks for AI researchers.
Public evaluation platform for agent long-term memory, defining a unified protocol with multi-track benchmarks and capability profiles to enable comparable leaderboard rankings.
Offers a Go-based AI gateway unifying OpenAI- and Anthropic-compatible APIs across many providers, with intelligent routing, failover, streaming, observability, guardrails, and cost tracking.
Desktop application for optimizing AI agents using production traces and reinforcement-style analysis. Imports browsing replays, ranks failure modes, and generates actionable fix suggestions for engineers improving hierarchical agent loops.
High-performance Rust-powered statusline HUD for Claude Code that displays current model, git branch, and usage information directly inside the developer's terminal workflow.
llama.cpp fork focused on local LLM inference efficiency, using KVarN, KV-cache precision tailoring, and low-bit quantization to support longer context with better precision in the same VRAM.
A large-scale realistic benchmark built from real-world vulnerabilities, including user-space and kernel cases, for evaluating AI agents' ability to develop working exploits.
An endpoint observability platform for security teams monitoring AI agent activity, combining on-device detection, optional pre-action blocking, and forensic session reconstruction to investigate automation behavior and policy violations.
A pure-Rust terminal workbench built with Zed gpui and Alacritty VT core, offering shells, persistent sessions, SSH access, and coding-agent collaboration for developers.
An enterprise AI API gateway unifying 30+ model providers behind OpenAI-compatible endpoints, with token controls, billing, failover and active-active clustering.
Supplies open-source sandbox infrastructure for AI workloads with a unified API over Firecracker, QEMU, and libkrun, supporting persistent microVMs, browser access, and file sharing.
An Electron-based desktop workspace for developers managing multiple pi Agent sessions across local project directories, with unified history browsing, session restore, Git integration, built-in terminal, model settings, and plugin support.
macOS menu-bar application for managing local ds4 model services with fast DeepSeek V4 Pro and Flash variants. It handles model downloads, service monitoring, and launching coding assistants with million-token context support.
A native macOS terminal workspace combining tabs, splits, a browser, and a file tree, with background AI agents and unified state management for collaborative command-line work.
Offers a terminal-based consciousness monitor for persistent-memory Hermes agents, visualizing memory, skills, projects, health status, and error-correction activity in real time for operators.
Creates a scalable benchmarking platform for multimodal Windows agents, providing authentic operating-system environments and large-scale parallel evaluation for researchers testing desktop automation.
Mobile-first web interface for managing multiple OpenCode AI agents from any device, offering session chat, file management, Git integration and task scheduling in a deployable PWA.
A single-file inference program for running the full Kimi K3 model on consumer hardware, plus an OpenAI-compatible local API server for chat and coding agents.
Local version-control and auditing tool for AI coding agents that records agent actions, attributes code changes to specific prompts, and lets developers inspect individual steps with surrounding context for debugging.
Implements an OpenAI-compatible gateway by reverse-engineering ChatGPT web protocols, supporting account pool management, text and GPT-Image models, plus batch image generation and editing.
Local-first dashboard for coding agents such as Claude Code and Codex, breaking each task into work receipts with tools used, files changed, tests run, time, and token costs.
Maintains a fork of the llama.cpp local LLM inference engine, supporting low-bit model formats and multiple hardware backends for developers building local inference services.
Local gateway and dashboard for Codex Desktop that handles account-pool routing, authentication, and task visualization for teams operating multiple Codex accounts.
An open-source LLM router and cost optimizer that automatically sends simple prompts to cheap or local models and complex ones to premium models via an OpenAI-compatible proxy.
Minimal modern terminal designed for AI coding workflows, offering sidebar workspaces, horizontal and vertical split panes, one-click agent launching, and live per-agent activity and workspace state.
Relay service with a mobile companion letting users remotely follow and control local Codex sessions from a phone while computation remains on the computer.
An operating layer for AI agent workloads combining an AI-native terminal, token-saving compression, eBPF observability, memory and sandboxed secure execution.
Helps TypeScript and Node.js developers inspect AI agent runs locally by turning manual steps, tool calls, model calls, logs, failures, and durations into readable terminal execution trees.
Native terminal built for parallel multi-agent workflows, organizing work through workspaces and sessions while exposing a full control API for dispatching coding tasks efficiently.
Protocol and gateway acting as a consequence firewall for machine actions, verifying exact authority before money, code, permission, or infrastructure changes and producing independently verifiable receipts.
An evaluation benchmark for general-purpose robot manipulation policies, combining simulated tasks with real-world validation procedures to help robotics researchers compare manipulation strategies reproducibly.
Local analytics dashboard for Claude Code users that visualizes session costs, model usage, project trends, and task-level details. Runs entirely offline without cloud services or telemetry for privacy-preserving usage analysis.
C++ ggml port of Nvidia's LocateAnything-3B open-vocabulary detection model enabling fast object localization on CPU and GPU for developers building vision applications.
An inference engine built on stock llama.cpp that runs oversized MoE models on memory-constrained phones by streaming only routed experts from flash without quality loss.
A benchmark series for context learning that evaluates reasoning and learning ability through workplace and everyday tasks with thousands of detailed annotation rules.
OpenCode plugin that routes calls through a Claude Max or Pro subscription via Meridian, automatically managing local proxy lifecycles for conflict-free multi-instance use.
A gateway service that converts authenticated Grok logins into endpoints compatible with multiple LLM API specifications, offering multi-worker processing and a web admin panel for operators.
vLLM patch with hand-written SM120 SASS kernels combining 2-bit MoE experts and an FP4 delta cache to recover precision, enabling frontier MoE models to run on consumer Blackwell GPUs.
Self-hosted web interface for remotely monitoring and controlling existing Claude and Codex CLI sessions from phones or other devices. Supports push notifications, tool approvals, file uploads, and device management without accounts or databases.
Containerized LLM inference toolboxes for developers on AMD Strix Halo hardware, tracking versioned releases and supporting both single-machine deployment and multi-node inference clusters.
Unified launch and management hub for multiple AI coding agents such as Claude Code, Codex CLI, and Gemini CLI, supporting parallel sessions across desktop, mobile, and messaging.
A hybrid LLM inference runtime for running very large models on consumer GPUs with limited VRAM. Uses GPU prefill, decoding, and hot-cold expert management to enable local execution for individual practitioners.
A command-line installer that fetches agent skills from npm packages and links them into multiple agent configuration directories, helping developers keep skill definitions synchronized across tools.
Offline evaluation toolkit that measures the maturity of coding-agent harnesses in seconds, scoring reliability practices and returning a prioritized remediation checklist for developers.
Local-first agent operating system that drives official Claude Code, Codex and OpenCode instances from browsers or chat apps while unifying workspace management on the user machine.
Local gateway for users converting Tabbit access into OpenAI-compatible APIs, bundling membership authentication and a browser extension for one-click cookie extraction to call Claude and GPT models.