π Performance
Benchmarks
Run batch evaluations on 10 science questions from the ingested Wikipedia corpus. Compare token usage, F1 score, and cost across all 3 pipelines.
Run Benchmark
checking TigerGraphβ¦Runs honest_benchmark.py --graph-native live against TigerGraph β real F1, LLM-judge, and BERTScore. No hardcoded or demo numbers.
10