deterministic-scoring

Here is 1 public repository matching this topic...

Basaltlabs-app / Gauntlet

Community-driven behavioral reliability benchmark for LLMs. 231 probes across 19 modules, deterministic scoring, perplexity correlation, layer sensitivity mapping, quant method capture, hardware-stratified community rankings. Every test contributes to the community dataset.

benchmark mcp community-driven model-evaluation ai-evaluation llm ollama sycophancy hallucination-detection llm-testing hardware-benchmark ai-trust trust-scoring behavioral-testing llm-benchmark deterministic-scoring

Updated Apr 17, 2026
Python

Improve this page

Add a description, image, and links to the deterministic-scoring topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the deterministic-scoring topic, visit your repo's landing page and select "manage topics."

Learn more

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

deterministic-scoring

Here is 1 public repository matching this topic...

Basaltlabs-app / Gauntlet

Improve this page

Add this topic to your repo