FazBrowse GitHub Viewer
|
Trending
|
URL:
|
Home
Tools:
[Download Repo ZIP]
[View Raw Code]
[Original HTTPS Page]
TensorSharp2/.github/workflows/test-matrix.yml at main · SciSharp/TensorSharp2 · GitHub
SciSharp
/
TensorSharp2
Public
forked from
zhongkaifu/TensorSharp
Notifications
You must be signed in to change notification settings
Fork
0
Star
6
Code
Pull requests
0
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Pull requests
Actions
Projects
Security and quality
Insights
Expand file tree
Breadcrumbs
TensorSharp2
/
.github
/
workflows
/
test-matrix.yml
Copy path
View runs
More file actions
More file actions
Latest commit
History
History
History
277 lines (258 loc) · 11.9 KB
Breadcrumbs
TensorSharp2
/
.github
/
workflows
/
test-matrix.yml
Copy path
File metadata and controls
277 lines (258 loc) · 11.9 KB
Raw
Copy raw file
Download raw file
Open symbols panel
Edit and raw actions
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
name
:
Engine Comparison (TensorSharp vs llama.cpp)
#
Cross-engine benchmark: TensorSharp vs llama.cpp on the same GGUF files.
#
#
This replaces the old TensorSharp.TestMatrix-driven matrix (which fanned out
#
over backends x features x env-var sweeps and regularly timed out). The
#
simplified workflow does exactly four things:
#
#
1. Clones and builds llama.cpp (CUDA) and sets it up for the harness.
#
2. Downloads the benchmark models from their Hugging Face pointers (the
#
`_hf` fields in benchmarks/engine_comparison/benchmark_config_ci.json)
#
and runs benchmarks/engine_comparison/run_matrix.py to compare
#
TensorSharp vs llama.cpp performance.
#
3. Compares output quality between the two engines (report.py's
#
"Output quality" section: cross-engine similarity of greedy outputs,
#
JSON validity, tool-call correctness).
#
4. Generates a single markdown report covering both performance and
#
quality (uploaded as an artifact + rendered into the job summary).
#
#
Two profiles:
#
- smoke (pull_request): one small model (gemma4-12b), four cheap scenarios
#
(text_short, function_call, json_mode, prefill_4k). The report is posted
#
as a PR comment. Superseded runs on the same PR are cancelled.
#
- full (workflow_dispatch / weekly schedule): the whole CI model +
#
scenario set; dispatch inputs can select a custom subset.
#
#
llama.cpp sources/build and the model files live in a persistent directory
#
on the runner ($HOME/tensorsharp-bench by default, overridable with the
#
BENCH_HOME repository variable), so only the first run on a fresh runner
#
pays the llama.cpp build + ~20 GB model download; repeat runs (including PR
#
smoke runs) only pay incremental costs.
#
#
Self-hosted runner required: [self-hosted, tensorsharp-cuda]
#
- Linux + NVIDIA GPU with the CUDA toolkit (nvcc) on PATH
#
- .NET 10 SDK, cmake, git, python3 + pip on PATH
#
- ~40 GB free disk for llama.cpp builds + model files
on
:
pull_request
:
branches
:
[main]
workflow_dispatch
:
inputs
:
models
:
description
:
"
Comma-separated model ids from benchmark_config_ci.json (empty = default set)
"
required
:
false
default
:
"
"
scenarios
:
description
:
"
Comma-separated scenario ids (empty = default set)
"
required
:
false
default
:
"
"
llama_ref
:
description
:
"
llama.cpp git ref to build (branch, tag or SHA)
"
required
:
false
default
:
"
master
"
schedule
:
#
Weekly regression point - Monday 03:00 UTC.
-
cron
:
"
0 3 * * 1
"
concurrency
:
#
One benchmark per ref. PR smoke runs are cancelled when new commits land;
#
full runs (dispatch / schedule) are long and expensive, so later triggers
#
queue behind them instead of cancelling.
group
:
engine-comparison-${{ github.ref }}
cancel-in-progress
:
${{ github.event_name == 'pull_request' }}
defaults
:
run
:
shell
:
bash
jobs
:
benchmark
:
name
:
Benchmark (tensorsharp-cuda)
runs-on
:
[self-hosted, tensorsharp-cuda]
#
PR smoke runs are small; only full runs need the long ceiling.
timeout-minutes
:
${{ github.event_name == 'pull_request' && 120 || 300 }}
env
:
#
pull_request -> trimmed smoke profile; dispatch/schedule -> inputs or
#
the full CI default set.
BENCH_PROFILE
:
${{ github.event_name == 'pull_request' && 'smoke (pull request)' || github.event_name == 'schedule' && 'full (weekly schedule)' || 'full / custom (manual dispatch)' }}
BENCH_MODELS
:
${{ github.event_name == 'pull_request' && 'gemma4-12b' || github.event.inputs.models || 'gemma4-12b,qwen36-35b-a3b' }}
BENCH_SCENARIOS
:
${{ github.event_name == 'pull_request' && 'text_short,function_call,json_mode,prefill_4k' || github.event.inputs.scenarios || 'text_short,text_long,multi_turn,function_call,json_mode,prefill_2k,prefill_4k,prefill_8k' }}
LLAMA_REF
:
${{ github.event.inputs.llama_ref || 'master' }}
steps
:
#
clean: false keeps the untracked incremental state that makes repeat
#
runs fast (TensorSharp.GGML.Native/build, ExternalProjects/ggml).
-
uses
:
actions/checkout@v4
with
:
submodules
:
recursive
clean
:
false
-
name
:
Resolve persistent benchmark directory
run
:
|
BENCH_HOME="${{ vars.BENCH_HOME }}"
BENCH_HOME="${BENCH_HOME:-$HOME/tensorsharp-bench}"
mkdir -p "$BENCH_HOME"
{
echo "BENCH_HOME=$BENCH_HOME"
echo "BENCH_MODEL_ROOT=$BENCH_HOME/models"
echo "BENCH_RESULTS=$GITHUB_WORKSPACE/bench-results"
} >> "$GITHUB_ENV"
{
echo "### Run profile"
echo ""
echo "- profile: \`$BENCH_PROFILE\`"
echo "- models: \`$BENCH_MODELS\`"
echo "- scenarios: \`$BENCH_SCENARIOS\`"
echo ""
} >> "$GITHUB_STEP_SUMMARY"
-
name
:
Show host info
run
:
|
uname -a
dotnet --version
cmake --version | head -1
python3 --version
nvidia-smi ||
true
df -h "$BENCH_HOME" .
-
name
:
Build TensorSharp native GGML library (CUDA)
run
:
bash TensorSharp.GGML.Native/build-linux.sh --cuda
-
name
:
Build TensorSharp.Server.Host
#
The benchmark harness launches the server by DLL, so build the host
#
application (TensorSharp.Server is the library it references).
run
:
dotnet build TensorSharp.Server.Host/TensorSharp.Server.Host.csproj -c Release
#
----- 1. Clone and build llama.cpp, set up its environment -----------
-
name
:
Clone / update llama.cpp
run
:
|
set -euo pipefail
LLAMA_SRC="$BENCH_HOME/llama.cpp"
if [ ! -d "$LLAMA_SRC/.git" ]; then
git clone https://github.com/ggml-org/llama.cpp "$LLAMA_SRC"
fi
git -C "$LLAMA_SRC" fetch --tags origin "$LLAMA_REF"
git -C "$LLAMA_SRC" checkout --detach FETCH_HEAD
echo "LLAMA_SRC=$LLAMA_SRC" >> "$GITHUB_ENV"
echo "llama.cpp @ $(git -C "$LLAMA_SRC" log -1 --format='%h %s')"
-
name
:
Build llama-server (CUDA)
run
:
|
set -euo pipefail
cmake -S "$LLAMA_SRC" -B "$LLAMA_SRC/build" \
-DCMAKE_BUILD_TYPE=Release \
-DGGML_CUDA=ON \
-DLLAMA_CURL=OFF \
-DLLAMA_BUILD_SERVER=ON
cmake --build "$LLAMA_SRC/build" --config Release --target llama-server -j "$(nproc)"
LLAMA_BIN="$LLAMA_SRC/build/bin/llama-server"
test -x "$LLAMA_BIN"
echo "BENCH_LLAMA_SERVER=$LLAMA_BIN" >> "$GITHUB_ENV"
{
echo "### Engine versions"
echo ""
echo "- TensorSharp: \`$(git rev-parse --short HEAD)\`"
echo "- llama.cpp: \`$(git -C "$LLAMA_SRC" rev-parse --short HEAD)\` (ref \`$LLAMA_REF\`)"
echo ""
} >> "$GITHUB_STEP_SUMMARY"
#
----- 2. Download models from their Hugging Face pointers ------------
-
name
:
Download benchmark models
env
:
HF_TOKEN
:
${{ secrets.HF_TOKEN }}
run
:
|
python3 -m pip install --user -q -U huggingface_hub requests
python3 benchmarks/engine_comparison/download_models.py \
--config benchmarks/engine_comparison/benchmark_config_ci.json \
--models "$BENCH_MODELS"
#
----- 2. Run the performance comparison ------------------------------
-
name
:
Run benchmark matrix (TensorSharp vs llama.cpp)
id
:
bench
working-directory
:
benchmarks/engine_comparison
run
:
|
set -euo pipefail
rm -rf "$BENCH_RESULTS"
python3 run_matrix.py --config benchmark_config_ci.json \
--models "$BENCH_MODELS" \
--scenarios "$BENCH_SCENARIOS"
#
----- 3 + 4. Quality comparison + combined report ---------------------
#
report.py aggregates the per-cell JSONs into one markdown report with
#
both the performance tables/ratios and the output-quality section.
#
These steps run even when the benchmark step itself failed partway
#
(some cells may still be reportable) but not when it never ran.
-
name
:
Generate performance + quality report
if
:
${{ !cancelled() && steps.bench.outcome != 'skipped' }}
working-directory
:
benchmarks/engine_comparison
run
:
|
python3 report.py --config benchmark_config_ci.json
cp ../../docs/engine_comparison_report.md "$GITHUB_WORKSPACE/engine-comparison-report.md"
-
name
:
Upload report
if
:
${{ !cancelled() && steps.bench.outcome != 'skipped' }}
uses
:
actions/upload-artifact@v4
with
:
name
:
engine-comparison-report
path
:
engine-comparison-report.md
retention-days
:
90
-
name
:
Upload per-cell results (JSON, logs, CSV)
if
:
${{ !cancelled() && steps.bench.outcome != 'skipped' }}
uses
:
actions/upload-artifact@v4
with
:
name
:
engine-comparison-results
path
:
bench-results
retention-days
:
30
-
name
:
Render report into job summary
if
:
${{ !cancelled() && steps.bench.outcome != 'skipped' }}
run
:
|
if [ -f engine-comparison-report.md ]; then
# GitHub job summaries cap at 1 MiB.
head -c 900000 engine-comparison-report.md >> "$GITHUB_STEP_SUMMARY"
else
echo "(no report produced)" >> "$GITHUB_STEP_SUMMARY"
fi
-
name
:
Fail on errored or empty benchmark cells
run
:
|
python3 - <<'EOF'
import json, os, sys
from pathlib import Path
results = Path(os.environ["BENCH_RESULTS"])
cells = [json.loads(p.read_text(encoding="utf-8"))
for p in results.glob("*.json")]
failed = [c for c in cells if c.get("status") == "fail"]
ok = [c for c in cells if c.get("status") == "ok"]
for c in failed:
print(f"FAIL {c['engine']}/{c['backend']}/{c['model']}/{c['scenario']}: "
f"{c.get('detail', '')[:300]}")
if failed:
sys.exit(f"{len(failed)} benchmark cell(s) failed - see bench-results logs")
if not ok:
sys.exit("no benchmark cell ran successfully - check model/binary setup")
print(f"{len(ok)} ok cells, 0 failed")
EOF
pr-comment
:
name
:
PR comment
needs
:
benchmark
if
:
github.event_name == 'pull_request' && always()
runs-on
:
ubuntu-latest
permissions
:
pull-requests
:
write
steps
:
-
uses
:
actions/download-artifact@v4
with
:
pattern
:
engine-comparison-report
merge-multiple
:
true
path
:
report
-
name
:
Post or update PR comment
uses
:
actions/github-script@v7
with
:
script
:
|
const fs = require('fs');
const marker = '<!-- engine-comparison-report -->';
let body = marker + '\n## Engine comparison — TensorSharp vs llama.cpp (PR smoke)\n\n';
const f = 'report/engine-comparison-report.md';
if (fs.existsSync(f)) {
const txt = fs.readFileSync(f, 'utf8');
// PR comment bodies cap at 65536 chars.
body += txt.length > 60000
? txt.slice(0, 60000) + '\n\n_(truncated — see the engine-comparison-report artifact for the full report)_'
: txt;
} else {
body += '_No report artifact was produced — the benchmark failed before generating results (see the workflow logs)._';
}
const { owner, repo } = context.repo;
const issue_number = context.issue.number;
const { data: comments } = await github.rest.issues.listComments({ owner, repo, issue_number });
const existing = comments.find(c => c.body && c.body.startsWith(marker));
if (existing) {
await github.rest.issues.updateComment({ owner, repo, comment_id: existing.id, body });
} else {
await github.rest.issues.createComment({ owner, repo, issue_number, body });
}
Back
|
FazBrowse Home
|
New Git URL