中文
Agent infrastructure

deepeval

confident-ai/deepeval

Open-source LLM evaluation framework similar to Pytest for unit-testing LLM applications, providing diverse metrics to assess models, prompts, and agent trajectory quality.