hpc-benchmark-agent
Python · SQLite · FastAPI · function calling
Ask a question in Chinese about PETSc benchmark runs; get an answer grounded in a database rather than in the model's memory.
-
Layered so the two halves stay independent.
-
A rule router as the baseline, not as a fallback.
-
Both routers return the same shape.
-
The eval set scores tool calls, not answer text.
-
Failure is explicit.
-
Reports are byte-deterministic.
-
Traceable and offline-testable.
Verify it in three minutes:
clone the repo, install, then run the tests — no API key needed.
PythonSQLiteFastAPI
OpenAI SDKpytestPETSc
MPISlurm
github.com/leyancode/hpc-benchmark-agent