| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
parent directory.. | ||||
Automated compatibility testing for the Robocode → Tank Royale API bridge. For every legacy robot jar in the LiteRumble collection it runs the same battle (robot vs. itself) on classic Robocode and on Tank Royale (robot wrapped via the bridge), then compares scores and errors.
| File | Purpose |
|---|---|
| compat_test.py | Orchestrator: staging, checkpointing, error logs, report generation |
| regression-set.json | The pinned watch list the regression gate measures against |
| parity-registry.json | Tracked, append-only observations for every jar or team |
| parity-registry.md | Generated review table for the parity registry |
| trace-robot/ | Robots that report their own state: TraceRobot for --trace, TurnSignProbe for settling a convention question by measurement |
| RcBattleWorker.java | Single-file worker (run uncompiled) driving the classic Robocode Control API |
| TrBattleWorker.java | Single-file worker (run uncompiled) driving the Tank Royale Battle Runner API |
Each robot test spawns fresh worker JVMs, so a hanging or crashing robot can never take down the harness — the orchestrator kills the whole process tree on timeout and records the failure.
JDK 17+ on PATH (java), Python 3.9+ (stdlib only).
A JDK 23 or older for the classic side. Classic Robocode installs a SecurityManager to sandbox robots, and JDK 24 removed SecurityManager support outright -- so -Djava.security.manager=allow is no longer a deprecation warning but a fatal VM error, and classic cannot start at all on a JDK 24+ default. The harness auto-detects a suitable JDK; override with --rc-java or COMPAT_RC_JAVA. The Tank Royale side is unaffected.
Robot collection at C:\Code\LiteRumble robots with roborumble, meleerumble, teamrumble subdirectories of .jar files.
Classic Robocode installation at C:\robocode; LiteRumble currently accepts client versions 1.10.3, 1.11.0, and 1.11.1, and the harness records and validates the installed version.
Built artifacts (all present after building the respective repos):
⚠️ The bot-api version matters: it must be protocol-compatible with the server embedded in the runner jar (an incompatible pairing leaves the robots idle, scoring 0 — newer Bot APIs fail loudly on this). Do not fall back to 0.33.1: its event queue drops deferred same-priority events, e.g. every other scan event for robots that call a blocking method such as fire() inside onScannedRobot.
All paths are defaults only — override with CLI flags (--collection-dir, --robocode-home, --runner-jar, --bridge-api-jar, --wrapper-jar, --bot-api-jar) or the corresponding COMPAT_* environment variables. Set COMPAT_DATA_DIR to redirect the checkpoint, report, error logs, and parity registry to an isolated directory; set COMPAT_WORK_DIR to isolate temporary engine and bot files.
Each division runs at its official rumble parameters, read from the classic installation's roborumble/ configuration. Constraint C-003 in the corpus: these are not ours to choose. A robot is tuned to its division and its ranking was earned at these settings, so measuring it at anything else produces a number describing behaviour the robot was never ranked on.
| Division | Battlefield | Rounds | Bots per battle |
|---|---|---|---|
| roborumble | 800x600 | 35 | 2 |
| meleerumble | 1000x1000 | 35 | 10 |
| teamrumble | 1200x1200 | 10 | teams |
--rounds overrides the round count for quick local runs and makes the result incomparable with the rumble, so the report records the setup each row was measured at.
Official melee cases use the fixed twelve-jar opponent pool in melee-opponents.json. Each subject runs with the first nine pool entries after excluding itself, keeping ten participants while avoiding accidental self-play. The setup recorded in the parity registry includes the pool and the selected opponent names and hashes.
cd compat-test
python compat_test.py # run the next 25-subject checkpoint
python compat_test.py --collections roborumble --limit 50
python compat_test.py --only Waylander # substring filter on jar name
python compat_test.py --rounds 5 # override the official round count
python compat_test.py --retry-failed # re-run only FAIL/ERROR robots
python compat_test.py --retry-unresolved # re-run every unresolved parity case
python compat_test.py --retest-cause lifecycle # re-run unresolved cases tagged to one cause
python compat_test.py --retest-cause lifecycle --repair 249cdaf # link retest to its repair
python compat_test.py --sync-registry --registry-manifest C:\path\to\run-manifest.json
python compat_test.py --force # re-run everything
python compat_test.py --report-only # just regenerate the reportpython compat_test.py --regression # re-measure the pinned watch list
python compat_test.py --regression --repeats 3 --band 20Re-measures every bot in regression-set.json over averaged repeats and reports movement from its recorded baseline. A single battle is not evidence about a bot -- one watched bot swings by a factor of forty between runs on classic alone -- so scores are averaged before they are judged, and the verdict is stated as movement from a baseline rather than as an absolute delta.
Every baseline in the shipped watch list is null, deliberately. The deltas recorded there were measured before the event-dispatch redesign and the Bot API upgrade, at ten rounds, in a setup matching no division exactly. They say why a bot is watched; they are not something to measure against. The gate reports NO BASELINE until a sweep at official parameters records real ones.
A bot marked noise is always reported and never fails the run. A gate that fires on the same bot every time is a gate people learn to ignore.
Fail-fast on a bridge-only exception. The classic side runs first, so its exception signatures are the baseline. When the Tank Royale side throws a signature classic did not produce, the battle is stopped there and the signature recorded -- no score, no averaging. A score difference is a quantity that repetition can resolve; the same bot throwing only under the bridge is a categorical fact that repetition cannot improve.
trace-robot/tracing/ holds robots that exist to answer a question the engines will not answer directly. They are not part of any sweep; each is run by hand, on both engines, when a disagreement needs settling.
TurnSignProbe is the worked example. It commands a known right turn, then prints the remaining turn from both paths a robot can read it through -- the peer's getter and the RobotStatus handed to onStatus -- so the two can be compared against each other and against classic.
# compile against the classic API, then run on each engine
javac -cp "$COMPAT_ROBOCODE_HOME/libs/*" -d work/probe-classes trace-robot/tracing/TurnSignProbe.java
python compat_test.py --conformance work/probe-classes --robot-class tracing.TurnSignProbe --engine rc --rounds 1
python compat_test.py --conformance work/probe-classes --robot-class tracing.TurnSignProbe --engine tr --rounds 1It settled AN-007. Classic reported +80 on both paths after a commanded right turn; the bridge reported −80 on the status path and +80 on the peer path, contradicting classic and itself. Reverting the fix reproduced the inversion and restoring it removed it, which is what turned an argument from documentation into a measurement.
The pattern generalises: when the two engines' conventions are in question, a robot compiled once against the classic API and run on both engines reports what each actually does. That is cheaper than reading two codebases and it cannot be wrong about the engines in the way a careful reading can.
python compat_test.py --trace # per-turn state from both engines
python compat_test.py --trace --trace-turns 80Compiles trace-robot/tracing/TraceRobot.java against the classic API, runs it on both engines, and prints the two per-turn streams side by side with the differing lines marked. This is the behavioural comparison the score-gap work needs; until now nothing could show behaviour at all, only the score at the end.
The robot reports from inside: neither engine will hand a per-turn view to the harness from outside, and instrumenting an engine would mean measuring a build no robot will ever run against. Reporting from inside measures what the robot perceives, which is the parity question exactly, and one class compiled against the classic API runs on both engines because reproducing that API is what the bridge is for.
Its movement is a fixed command sequence rather than a strategy, so that what differs between the two traces is the engine rather than the robot.
Expect divergence to grow once the robots interact -- Tank Royale has no seed, so the two battles are not the same battle. The early turns, before anything is scanned, are where a real mapping fault shows.
First observation from this tool, recorded and not yet explained: classic's first traced turn is turn=0 and the bridge's is turn=1, and the bridge's state at its turn N matches classic's at turn N. So the robot's first execution appears to happen a turn later under the bridge rather than the turn number being offset. Worth noting alongside it that BotPeer.getTime() returns the Tank Royale turn number unchanged while the status mapper rebases the round number by one -- the codebase already treats the two engines' counting conventions as differing, in one place and not the other. Needs investigation before anything is concluded; it belongs to the score-gap milestones.
Progress is written to test_progress.json after every robot (atomic replace). Interrupt at any time (Ctrl+C, crash, reboot) — re-running resumes from the first untested robot. Completed robots are never re-run unless --force (everything) or --retry-failed (failures only) is given.
Full error details land in errors/robocode/<robot>.log and errors/tank-royale/<robot>.log; the master table is compatibility_report.md.
parity-registry.json is the tracked evidence carrier. It appends each subject observation with the exact jar identity, setup, engine artifacts, normalized errors, and focused retest link. Normalized error origins skip engine implementation frames from both classic and the bridge so the first legacy application frame remains comparable across engines. It also retains an append-only diagnosis history, so a later triage decision cannot rewrite an earlier one. parity-registry.md renders the current status of every subject for review. Import an existing checkpoint with --sync-registry; after a diagnosis, tag a case with --set-cause <subject> <cause> <owner> and rerun that cause with --retest-cause <cause> --repair <commit-or-PR>.
Pass --capture-skipped-turns to enable opt-in records for Tank Royale bridge callbacks; capture is off by default. Each registry observation stores tank_royale.skipped_turn_telemetry with a status and, only when capture is complete, an events list of unique {bot_id, round, turn} values sorted by round and turn, including warm-up. disabled, unavailable, and incomplete use events: null; captured with events: [] means the completed run delivered no skipped-turn callbacks to the bridge. Per-bot readiness and completion-count markers let the harness detect an older bridge jar or a truncated capture. The data describes callbacks delivered to the bridge and does not claim to record server detections that never reach it.
--turn-timeout-micros overrides Tank Royale's per-turn timeout for a local probe. A large value makes a completed no-event capture practical to verify without depending on a bot narrowly meeting the default turn deadline.
Use the conformance probe to check a forced-skip run locally:
python compat_test.py --conformance conformance-robots --conformance-source conformance-robots/conformance/probes/SkippedTurnProbe.java --robot-class conformance.probes.SkippedTurnProbe --engine tr --rounds 1 --capture-skipped-turnsThe conformance JSON includes skipped_turn_telemetry at the top level. On a parity sweep, use --capture-skipped-turns and inspect the per-observation field in parity-registry.json. Repeated --confirm-score runs keep one record per attempt under confirmation.skipped_turn_telemetry_runs; the regression gate prints the same per-attempt JSON to stdout.
Every row states the setup it was measured at. The report is regenerated from the state file long after the battles ran, so a single header describing the current configuration would restate every stored row as though it had been measured under today's settings -- which is how a stale result stops looking stale. A row measured at its division's official parameters reads official; one measured at anything else says what it was measured at; one stored before the harness recorded per-row setups reads unrecorded. --report-only is therefore safe to run over a state file holding results from several eras.
| Status | Meaning |
|---|---|
| PASS | Both ran; |score delta| ≤ threshold (default 25%); no TR-only errors |
| DISCREPANCY (score) | Both ran, but scores diverge beyond the threshold |
| DISCREPANCY (errors) | TR side threw errors the RC side didn't |
| FAIL (TR) / FAIL (RC) / FAIL (both) | The battle did not complete on that side |
| SKIPPED-TR | Team jar: classic result recorded, TR skipped (wrapper has no team support yet) |
| Back | FazBrowse Home | New Git URL |