Cross-platform desktop assistant for managing configurations, model switching, and sessions across Claude Code, Codex, OpenCode, and other coding agents centrally.
A resident CLI proxy that filters and compresses output from common development commands to cut token usage in AI coding sessions, packaged as a dependency-free Rust binary.
Free open-source AI gateway unifying access to over 1,200 models from 352 providers, offering quota-aware automatic fallback and token compression for coding agents.
Gateway utility for coding-agent users seeking free token access and unified model management, adding multi-provider failover across Claude Code, Codex, Pi, OpenCode, and other supported harnesses.
Open-source inference engine for running LLMs, vision, audio, and image models locally on diverse hardware, exposing a unified API compatible with popular model-serving interfaces.
Open-source gateway consolidating Claude, OpenAI, Gemini, and Grok subscriptions into one proxy endpoint, supporting shared usage and seamless integration with native developer tools.
A terminal-resident multiplexer for managing multiple AI agent sessions, preserving state across disconnects while supporting remote aggregation and plugin extensions for developers.
A persistent terminal runtime for coding agents that maintains connections after disconnects, aggregates multiple machines, and exposes agent-native controls for long-running development tasks.
Pure-C inference engine with zero dependencies for running very large mixture-of-experts models on commodity hardware by streaming experts from disk across unified memory and storage layers.
High-performance serving framework for engineers deploying large language and multimodal models, providing optimized inference, scalable deployment, and production-oriented runtime features.
Low-memory inference library using layer-wise offloading to run 70B to 671B large language models for inference on a single GPU with limited VRAM, aimed at researchers and individual developers.
Open tooling for the agent-skills ecosystem that fetches reusable skills from multiple Git sources via npx and injects them into dozens of coding agents.
Acts as a routing gateway connecting coding tools such as Claude Code, Cursor, and Copilot to more than forty model providers, with automatic fallback and token-saving optimizations.
A unified gateway aggregating dozens of free large language models behind one OpenAI-compatible endpoint, providing smart routing, failover, and encrypted key management for experimentation.
Adds a status display to Claude Code showing context usage, active tools, running subagents, and to-do progress so developers can monitor ongoing work in real time.
Open-source AI engineering platform covering experiment tracking, model registry, evaluation, monitoring, observability, and gateway services across agent and LLM application lifecycles.
An open-source macOS terminal based on Ghostty that provides vertical tabs, multi-session management, and notifications to help AI coding agents and developers handle parallel tasks.
Packages LLM weights together with an inference engine into single-file executables. Users run large models locally across operating systems without installation, including speech transcription tools.
An evaluation and red-teaming toolkit for LLM applications that supports multi-model comparison, vulnerability scanning, and CI/CD integration for AI quality and security testing.
Open-source desktop application and team control plane for sharing skills and MCP capabilities across agents such as Codex for collaborative development workflows.
Remote-control service for centrally managing multiple local coding agents from mobile, web, and desktop clients, aimed at developers who operate agents across devices.
A high-performance large language model deployment engine based on machine learning compilation, enabling developers to run LLMs efficiently across diverse chips and platforms through a unified API.
Delivers a local inference engine for DeepSeek 4 Flash and PRO models on consumer hardware, supporting Metal, CUDA, and ROCm backends with multi-GPU parallel serving for developers and researchers.
Reference stack for securely running AI agents such as Hermes and OpenClaw inside NVIDIA OpenShell, providing managed inference and lifecycle controls for deployments.
Open-source observability platform for LLM applications, RAG pipelines, and agents, offering tracing, evaluation, prompt optimization, and production monitoring for AI engineering teams.
Local LLM inference server optimized for Apple Silicon, featuring continuous batching and memory-plus-SSD tiered caching with convenient menu-bar management for Mac users.
Open-source observability platform unifying logs, metrics, traces, and frontend monitoring in a single binary deployment, offering teams a lower-cost alternative to Datadog and Elasticsearch.
Stays in the menu bar to aggregate AI coding usage, quotas, and reset countdowns across OpenAI Codex and Claude Code without requiring login, aimed at developers tracking limits.
Evaluation framework and open benchmark registry for large language models and systems, allowing researchers to create custom tests and run standardized assessments of model capabilities.
High-performance browser-side LLM inference engine built on WebGPU and compatible with OpenAI APIs, distributed as npm packages for developers building web-based AI applications.
Open-source LLM evaluation framework similar to Pytest for unit-testing LLM applications, providing diverse metrics to assess models, prompts, and agent trajectory quality.
Desktop tool centrally managing accounts, quotas, multi-instance sessions, and wake-up automation across AI coding IDEs including Codex, Copilot, Cursor, and Gemini CLI.
A security scanner for AI agent skills that detects vulnerabilities, prompt injection, data exfiltration, and supply-chain risks, producing scored reports before installation for developers.
Browser extension for developers using Codex that securely exports a logged-in ChatGPT session locally to generate a spec-compliant auth.json backup file.
Comprehensive observability, tracing, evaluation, and monitoring platform for LLM and multi-agent applications, including self-hosted dashboard and execution analysis views for developers.
A universal provider proxy that lets Codex CLI, apps, SDKs, and Claude Code run on arbitrary models including Claude, Gemini, Grok, DeepSeek, and Ollama by translating streaming and tool calls while managing account pools.
Evaluation toolkit for LLM and RAG applications offering objective metrics, synthetic test data generation, and production feedback analysis to help developers measure quality and improve reliability.
Evaluates AI agents through independent audits combining dozens of checks across five pillars, quickly scoring whether an agent behaves as intended for human or automated review.
Gateway bridging local AI coding assistants such as Claude Code, Cursor, and Codex to messaging platforms including Feishu, Slack, Telegram, and Discord without requiring public IP.
Lightweight LLM inference engine implemented from scratch in about 1200 lines of Python, compatible with vLLM interfaces and offering prefix caching, tensor parallelism, and fast offline inference.
Secure, fast, and scalable isolated code execution sandbox runtime for AI agents, supporting multi-language SDKs and container orchestration as a general sandbox platform.
Cloud infrastructure providing secure isolated sandboxes for AI agents. Executes code and controls browsers or desktops, with self-hosting options for developers building agent tooling.
Open-source second browser built for AI agents, offering isolated tabs, real-time dashboards, and session replay within a dedicated runtime environment.
Open-source high-speed AI gateway providing a unified API for over 1,600 LLMs, offering routing, load balancing, retries, fallbacks, and 50+ guardrails for developers managing LLM traffic.
Highly customizable statusline for the Claude Code CLI, displaying model usage, session context, and version information with powerline support and multiple themes.
High-performance secure sandbox service for AI agents, supporting single-machine and multi-node scaling with millisecond-level hardware-isolated startup for concurrent execution workloads.
LLM serving tool for deploying open models such as Llama and DeepSeek. Exposes OpenAI-compatible APIs locally or in the cloud for developers running inference services.
MLOps platform and Python library for machine learning and large models, providing experiment tracking, visualization, evaluation, and full lifecycle model management.
Local web dashboard for the Hermes agent runtime, providing multi-agent chat, session management, scheduled jobs, usage analytics, and cross-device coordination.
Open-source inference serving system for deploying models from multiple AI frameworks across cloud and edge, supporting concurrent batching and streaming for ML platform engineers.
Developer inspection and debugging tool for MCP servers, offering web, command-line, and terminal interfaces to test, validate, and troubleshoot Model Context Protocol services.
A launcher enabling the Codex app to use ChatGPT Web models, including Pro accounts, with account and model detection, installation setup, streaming, context, tools, and image support.
Python bindings for llama.cpp enabling local large language model inference, completion APIs, and OpenAI-compatible server deployment for developers running models on personal hardware.
Implements a WebAssembly-based operating system runtime for AI agents, offering sandboxed isolation, security controls, and portable state management for agent workloads.
SDK providing identity, communication and payment layers for AI agents, letting developers wrap agents from any framework with encrypted identity plus A2A messaging and USDC billing.
One-stop open-source model serving library deploying language, audio, and multimodal models with a single command, exposing unified production-grade inference APIs for developers.
Offers a generative AI red-teaming test suite that probes diverse large language models for hallucinations, prompt injection, jailbreaks, and other vulnerabilities to assess overall model safety.
Gives AI agents a persistent computer workspace with containerized and isolated execution backends plus a unified entry point for developers running agent tasks.
A model-serving framework for building high-performance inference APIs from arbitrary AI models, offering containerized deployment and batch-inference optimization for production use.
Secure private sandbox runtime for autonomous agents, enforcing declarative policies to isolate and control file, network and credential access during tool execution.
Manages multiple parallel coding-agent sessions inside a single terminal interface, using tmux and Git worktree isolation to enable concurrent development and streamlined change review.
Open agent operating system community edition providing command-line runtime, isolated environments, composable capsules, and model connectivity for building agent applications.
Supplies an OpenAI-compatible open-source agent API service supporting diverse models and infrastructure, offering conversation handling, vector capabilities, and orchestration for building conversational agents.
Provides an on-device large model inference runtime for Snapdragon platforms, supporting NPU, GPU, and CPU execution through cross-platform SDKs and OpenAI-like service interfaces.
Local-first lightweight microVM runtime and embeddable SDK that provides millisecond-start hardware-isolated execution for untrusted workloads, designed for developers building AI agents and code-execution tools.
Inference engine written in portable C99 runs trillion-parameter Kimi K3 models on a single CPU with about 8GB RAM, using disk streaming and no BLAS, framework, or GPU dependencies.
Toolkit for compressing, quantizing, deploying, and serving large language models with high-performance inference, supporting multiple engines and online API serving for production use.
Provides an open-source observability framework for ML and LLM systems with over 100 metrics. It helps engineers evaluate, test, and monitor tabular models, generative AI applications, and data pipelines.
Multi-account API gateway for developers using Grok, aggregating Grok Build, Web, and Console channels into one unified interface for model access and management.
Acts as a skill management layer for AI agents, helping developers publish, retrieve, evaluate, share, and continuously evolve skills across different agent frameworks.
High-performance LLM inference engine written in Rust, supporting automatic loading of multiple models, quantization and multimodal inference through an OpenAI-compatible serving interface.
Open evaluation platform for large language models that benchmarks diverse models across hundreds of datasets to enable systematic performance comparison.
Open-source OpenTelemetry-based observability extension for LLM applications, providing instrumentation for popular models and vector databases that integrates with existing monitoring stacks.
Acts as a local proxy for Claude Code that renders text context as images to reduce token usage, offering dashboard visibility and a warm mode for developers managing large contexts.
Community registry service for MCP server discovery and distribution, functioning like an app store where developers can publish, query and manage servers for MCP clients.
Offers an open-source desktop workspace for collaborating with AI agents through multi-session management, custom API and MCP integrations, plus a built-in documentation hub for human-machine workflows.
Proxy and data plane for agentic applications that centralizes agent routing, model routing, observability tracing, and safety guardrails, aimed at platform teams operating production AI agent systems.
Efficient inference-serving framework for omni-modal models, unifying text, image, audio, video, and action inputs under one serving architecture for developers deploying multimodal applications.
One-command management tool that scans a project's tech stack, selects matching AI agent skills, verifies compatibility, and installs the complete skill set automatically.
An efficient Gemma 4 inference runtime for Apple-silicon Macs, streaming experts under low-memory budgets and bundling a demo app with evaluation tooling.
LLM routing gateway for autonomous agents such as OpenClaw, aggregating multiple providers with local millisecond-level model selection and supporting autonomous USDC payments for agent-driven API usage.
A governance toolkit for autonomous AI agents offering policy enforcement, zero-trust identity, execution sandboxing, and auditability. Helps platform teams mitigate reliability risks across the OWASP Agentic Top 10.
A portable lightweight self-contained micro-VM manager that lets developers launch isolated environments quickly, package them as single files, and run branched instances locally.
A pixel-style desktop pet that monitors coding agents such as Claude Code, Codex, and Cursor, visualizing their running status and progress so developers can glance at activity.
Rebuilds the Legado reading app interface with Material Design 3, providing a modernized cross-device experience for managing and reading local and online books.
Packs browser, shell, file access, MCP services, and VSCode Server into one Docker container, providing developers a secure isolated sandbox for running AI agents.
Provides local-first search, analytics, and token-usage statistics for coding-agent sessions across Claude Code, Codex, and more than twenty other agents for cost tracking.
Routes requests intelligently across mixed-model fleets spanning cloud, data center, and edge, helping platform engineers schedule inference efficiently at system level.
Open-source authentication gateway that connects AI agents to more than 1,500 SaaS providers for authorized actions through SDK, CLI, MCP, HTTP, and OpenAPI interfaces.
Serves optimized large language models directly from users own GPUs and NPUs, letting developers run local AI apps privately through standard inference interfaces.
A framework for evaluating and improving agents, supporting parallel multi-agent benchmarks, custom environments, and training-data generation for reinforcement learning workflows.
Command-line tool tracking token usage and costs across many AI coding agents, with global leaderboards plus two-dimensional and three-dimensional contribution visualizations.
Enforces OS-level filesystem and network restrictions on arbitrary processes without containers, offering lightweight sandboxing for safely running AI agents and untrusted code.