| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
Node Scraper is a tool which performs automated data collection and analysis for the purposes of system debug. For details on what data is collected and analyzed, see the plugin reference table.
Node Scraper is published on PyPI as amd-node-scraper. Install it with Python 3.9 or newer:
pip install amd-node-scraperWith uv:
uv pip install amd-node-scraperUse a virtual environment if you prefer. After installation, confirm the CLI is available:
node-scraper --helpNode Scraper requires Python 3.9+ for installation. After cloning this repository, run dev-setup.sh with source. This script uses uv to create a Python 3.9+ virtual environment, perform an editable install, and configure pre-commit hooks.
source dev-setup.shAlternatively, follow these manual steps:
curl -LsSf https://astral.sh/uv/install.sh | shuv venv venv --python 3.9
source venv/bin/activateOn Debian/Ubuntu without uv, you can use python3.9 -m venv venv instead.
uv pip install -e ".[dev]"This installs Node Scraper in editable mode with development dependencies. To verify: node-scraper --help
Equivalent using pip:
python3 -m pip install --editable .[dev] --upgradepre-commit installSets up pre-commit hooks for code quality checks. On Debian/Ubuntu, you may need: sudo apt install pre-commit
The Node Scraper CLI can be used to run Node Scraper plugins on a target system. The following CLI options are available:
usage: cli.py [-h] [--version] [--sys-name STRING]
[--sys-location {LOCAL,REMOTE}]
[--sys-interaction-level {PASSIVE,INTERACTIVE,DISRUPTIVE}]
[--sys-sku STRING] [--sys-platform STRING]
[--plugin-configs LIST] [--system-config STRING]
[--connection-config STRING] [--log-path STRING]
[--log-level {CRITICAL,FATAL,ERROR,WARN,WARNING,INFO,DEBUG,NOTSET}]
[--no-console-log] [--gen-reference-config] [--skip-sudo]
{summary,run-plugins,describe,gen-plugin-config,compare-runs,show-redfish-oem-allowable}
...
node scraper CLI
positional arguments:
{summary,run-plugins,describe,gen-plugin-config,compare-runs,show-redfish-oem-allowable}
Subcommands
summary Generates summary csv file
run-plugins Run a series of plugins
describe Display details on a built-in config or plugin
gen-plugin-config Generate a config for a plugin or list of plugins
compare-runs Compare datamodels from two run log directories
show-redfish-oem-allowable
Fetch OEM diagnostic allowable types from Redfish
LogService (for oem_diagnostic_types_allowable)
options:
-h, --help show this help message and exit
--version show program's version number and exit
--sys-name STRING System name (default: <current hostname>)
--sys-location {LOCAL,REMOTE}
Location of target system (default: LOCAL)
--sys-interaction-level {PASSIVE,INTERACTIVE,DISRUPTIVE}
Specify system interaction level, used to determine
the type of actions that plugins can perform (default:
INTERACTIVE)
--sys-sku STRING Manually specify SKU of system (default: None)
--sys-platform STRING
Specify system platform (default: None)
--plugin-configs LIST
Comma-separated built-in names and/or plugin config
JSON paths (e.g. --plugin-
configs=NodeStatus,/path/c.json). Built-ins:
AllIbPlugins, NodeStatus (default: None)
--system-config STRING
Path to system config json (default: None)
--connection-config STRING
Path to connection config json (default: None)
--log-path STRING Specifies local path for node scraper logs, use 'None'
to disable logging (default: .)
--log-level {CRITICAL,FATAL,ERROR,WARN,WARNING,INFO,DEBUG,NOTSET}
Change python log level (default: INFO)
--no-console-log Write logs only to nodescraper.log under the run
directory; do not print to stdout. If no run log
directory would be created (e.g. --log-path None),
uses ./scraper_logs_<host>_<timestamp>/ like the
default layout. (default: False)
--gen-reference-config
Generate reference config from system. Writes to
./reference_config.json. (default: False)
--skip-sudo Skip plugins that require sudo permissions (default:
False)Node Scraper can operate in two modes: LOCAL and REMOTE, determined by the --sys-location argument.
To use remote execution, specify --sys-location REMOTE and provide a connection configuration file with --connection-config.
node-scraper --sys-name <remote_host> --sys-location REMOTE --connection-config ./connection_config.json run-plugins DmesgPluginIn-band (SSH) connection:
{
"InBandConnectionManager": {
"hostname": "remote_host.example.com",
"port": 22,
"username": "myuser",
"password": "mypassword",
"key_filename": "/path/to/private/key"
}
}A sample is in config/connection-config_inband.example.json. Use password or key_filename.
Redfish (BMC) connection for Redfish-only plugins:
{
"RedfishConnectionManager": {
"host": "bmc.example.com",
"port": 443,
"username": "admin",
"password": "secret",
"use_https": true,
"verify_ssl": true,
"api_root": "redfish/v1"
}
}OOB SSH plugins use this same single-host RedfishConnectionManager block and open SSH to that BMC. A sample is in config/connection-config_oob.example.json.
Redfish plugins can collect from multiple BMCs concurrently. In-band plugins and OOB SSH plugins stay single-host. If this config has no top-level host, OOB SSH plugins are skipped. Add a top-level host when those plugins should still run against one BMC.
{
"RedfishConnectionManager": {
"targets": [
{
"target_key": "node-a",
"host": "bmc-node-a.example.com",
"username": "admin",
"password": "secret",
"use_https": true,
"verify_ssl": false,
"timeout_seconds": 30
},
{
"target_key": "node-b",
"host": "bmc-node-b.example.com",
"username": "admin",
"password": "secret",
"use_https": true,
"verify_ssl": false,
"timeout_seconds": 30
}
],
"max_workers": 32
}
}Multi-target mode applies to Redfish plugins (RedfishEndpointPlugin, RedfishOemDiagPlugin, and other plugins based on OOBandDataPlugin). Targets are collected concurrently. Wall-clock time follows the slowest target.
A target that fails to connect or collect does not fail the run when another target succeeds. That plugin result is a warning, and analysis still runs for the targets that returned data. The run fails when every target fails, or when analysis of collected data reports an error.
Per-target results are written to <plugin>/<collector>/<target_key>/. Single-target results stay in <plugin>/<collector>/. The same layout is used for the analyzer directory and for the AMC SSH-proxy plugin.
A ready-to-edit sample is in config/connection-config_redfish_multi_target.example.json. Replace the example hosts and password, then pass it with --connection-config.
Notes:
Plugins to run can be specified in two ways, using a plugin JSON config file or using the 'run-plugins' sub command. These two options are not mutually exclusive and can be used together.
You can use the describe subcommand to display details about built-in configs or plugins. List all built-in configs:
node-scraper describe configShow details for a specific built-in config
node-scraper describe config <config-name>List all available plugins**
node-scraper describe pluginShow details for a specific plugin
node-scraper describe plugin <plugin-name>The plugins to run and their associated arguments can also be specified directly on the CLI using the 'run-plugins' sub-command. Using this sub-command you can specify a plugin name followed by the arguments for that particular plugin. Multiple plugins can be specified at once.
You can view the available arguments for a particular plugin by running node-scraper run-plugins <plugin-name> -h:
usage: node-scraper run-plugins BiosPlugin [-h] [--collection {True,False}] [--analysis {True,False}] [--system-interaction-level STRING]
[--data STRING] [--exp-bios-version [STRING ...]] [--regex-match {True,False}]
options:
-h, --help show this help message and exit
--collection {True,False}
--analysis {True,False}
--system-interaction-level STRING
--data STRING
--exp-bios-version [STRING ...]
--regex-match {True,False}
Examples
Run a single plugin
node-scraper run-plugins BiosPlugin --exp-bios-version TestBios123Run multiple plugins
node-scraper run-plugins BiosPlugin --exp-bios-version TestBios123 RocmPlugin --exp-rocm TestRocm123Run plugins without specifying args (plugin defaults will be used)
node-scraper run-plugins BiosPlugin RocmPluginUse plugin configs and 'run-plugins'
node-scraper run-plugins BiosPluginThe 'gen-plugin-config' sub command can be used to generate a plugin config JSON file for a plugin or list of plugins that can then be customized. Plugin arguments which have default values will be prepopulated in the JSON file, arguments without default values will have a value of 'null'.
Examples
Generate a config for the DmesgPlugin:
node-scraper gen-plugin-config --plugins DmesgPluginThis would produce the following config:
{
"global_args": {},
"plugins": {
"DmesgPlugin": {
"collection": true,
"analysis": true,
"system_interaction_level": "INTERACTIVE",
"data": null,
"analysis_args": {
"analysis_range_start": null,
"analysis_range_end": null,
"check_unknown_dmesg_errors": true,
"exclude_category": null,
"interval_to_collapse_event": 60,
"num_timestamps": 3
}
}
},
"result_collators": {}
}Running DmesgPlugin with a dmesg log file:
Instead of collecting dmesg from the system, you can analyze a pre-existing dmesg log file using the --data argument:
node-scraper --run-plugins DmesgPlugin --data /path/to/dmesg.log --collection FalseThis will skip the collection phase and directly analyze the provided dmesg.log file.
Custom Error Regex Example:
You can extend the built-in error detection with custom regex patterns. Create a config file with custom error patterns:
{
"global_args": {},
"plugins": {
"DmesgPlugin": {
"analysis_args": {
"check_unknown_dmesg_errors": false,
"interval_to_collapse_event": 60,
"num_timestamps": 3,
"error_regex": [
{
"regex": "MY_CUSTOM_ERROR.*",
"message": "My Custom Error Detected",
"event_category": "SW_DRIVER",
"event_priority": 3
},
{
"regex": "APPLICATION_CRASH: .*",
"message": "Application Crash",
"event_category": "SW_DRIVER",
"event_priority": 4
}
],
"priority_override_rules": [
{
"message": "Application Crash",
"new_priority": "ERROR"
},
{
"event_category": "SW_DRIVER",
"new_priority": "WARNING"
}
]
}
}
},
"result_collators": {}
}Save this to dmesg_custom_config.json and run:
node-scraper --plugin-configs=dmesg_custom_config.json run-plugins DmesgPluginThe compare-runs subcommand compares datamodels from two run log directories (e.g. two nodescraper_log_* folders). By default, all plugins with data in both runs are compared.
Basic usage:
node-scraper compare-runs <path1> <path2>Exclude specific plugins from the comparison with --skip-plugins:
node-scraper compare-runs path1 path2 --skip-plugins SomePluginCompare only certain plugins with --include-plugins:
node-scraper compare-runs path1 path2 --include-plugins DmesgPluginShow full diff output (no truncation of the Message column or limit on number of errors) with --dont-truncate:
node-scraper compare-runs path1 path2 --include-plugins DmesgPlugin --dont-truncateYou can pass multiple plugin names to --skip-plugins or --include-plugins.
The show-redfish-oem-allowable subcommand fetches the list of OEM diagnostic types supported by your BMC (from the Redfish LogService OEMDiagnosticDataType@Redfish.AllowableValues). Use it to discover which types you can put in oem_diagnostic_types_allowable and oem_diagnostic_types in the Redfish OEM diag plugin config.
Requirements: A Redfish connection config (same as for RedfishOemDiagPlugin).
Command:
node-scraper --connection-config connection-config.json show-redfish-oem-allowable --log-service-path "redfish/v1/Systems/UBB/LogServices/DiagLogs"Output is a JSON array of allowable type names (e.g. ["Dmesg", "JournalControl", "AllLogs", ...]). Copy that list into your plugin config’s oem_diagnostic_types_allowable if you want to match your BMC.
Redfish OEM diag plugin config example
Use a plugin config that points at your LogService and lists the types to collect. Logs are written under the run log path (see --log-path).
{
"name": "Redfish OEM diagnostic logs",
"desc": "Collect OEM diagnostic logs from Redfish LogService. Requires Redfish connection config.",
"global_args": {},
"plugins": {
"RedfishOemDiagPlugin": {
"collection_args": {
"log_service_path": "redfish/v1/Systems/UBB/LogServices/DiagLogs",
"oem_diagnostic_types_allowable": [
"JournalControl",
...
"AllLogs",
],
"oem_diagnostic_types": ["JournalControl", "AllLogs"],
"task_timeout_s": 600
},
"analysis_args": {
"require_all_success": false
}
}
},
"result_collators": {}
}How to use
node-scraper --connection-config connection-config.json --plugin-config plugin_config_redfish_oem_diag.json run-plugins RedfishOemDiagPluginThe RedfishEndpointPlugin collects Redfish URIs (GET responses) and optionally runs checks on the returned JSON. It requires a Redfish connection config (same as RedfishOemDiagPlugin).
Multi-target support: RedfishEndpointPlugin collects from each BMC in the targets list at the same time. Use the Redfish multi-target connection config. The same uris and checks apply to every target. Per-target results are written to redfish_endpoint_plugin/redfish_endpoint_collector/<target_key>/ under the run log directory. A BMC that cannot be reached is reported as a warning when another target succeeds.
How to run
node-scraper --connection-config connection-config.json --plugin-config plugin_config_redfish_endpoint.json run-plugins RedfishEndpointPluginSample plugin config (plugin_config_redfish_endpoint.json):
{
"name": "RedfishEndpointPlugin",
"desc": "Redfish endpoint: collect URIs and optional checks",
"global_args": {},
"plugins": {
"RedfishEndpointPlugin": {
"collection_args": {
"uris": [
"/redfish/v1/",
"/redfish/v1/Systems/1",
"/redfish/v1/Chassis/1/Power"
]
},
"analysis_args": {
"checks": {
"/redfish/v1/Systems/1": {
"PowerState": "On",
"Status/Health": { "anyOf": ["OK", "Warning"] }
},
"/redfish/v1/Chassis/1/Power": {
"PowerControl/0/PowerConsumedWatts": { "max": 1000 }
}
}
}
}
},
"result_collators": {}
}collection_args
analysis_args
The 'summary' subcommand can be used to combine results from multiple runs of node-scraper to a single summary.csv file. Sample run:
node-scraper summary --search-path /<path_to_node-scraper_logs>This will generate a new file '/<path_to_node-scraper_logs>/summary.csv' file. This file will contain the results from all 'nodescraper.csv' files from '/<path_to_node-scarper_logs>'.
A plugin JSON config should follow the structure of the plugin config model defined here. The globals field is a dictionary of global key-value pairs; values in globals will be passed to any plugin that supports the corresponding key. The plugins field should be a dictionary mapping plugin names to sub-dictionaries of plugin arguments. Lastly, the result_collators attribute is used to define result collator classes that will be run on the plugin results. By default, the CLI adds the TableSummary result collator, which prints a summary of each plugin’s results in a tabular format to the console.
{
"globals_args": {},
"plugins": {
"BiosPlugin": {
"analysis_args": {
"exp_bios_version": "TestBios123"
}
},
"RocmPlugin": {
"analysis_args": {
"exp_rocm_version": "TestRocm123"
}
}
}
}Global args can be used to skip sudo plugins or enable/disble either collection or analysis. Below is an example that skips sudo requiring plugins and disables analysis.
"global_args": {
"collection_args": {
"skip_sudo" : 1
},
"collection" : 1,
"analysis" : 0
},A plugin config can be used to compare the system data against the config specifications. Built-in configs include NodeStatus (a subset of plugins) and AllIbPlugins (runs every registered in-band plugin with default arguments—useful for generating a reference config from the full system).
NodeStatus plus additional plugins — built-in configs merge with plugins named after run-plugins. Values are comma-separated; pass as --plugin-configs=… or --plugin-configs … (same as other optional flags), e.g. --plugin-configs=NodeStatus,/path/extra.json. Examples:
node-scraper --plugin-configs=NodeStatus run-plugins PciePluginnode-scraper --log-path ./logs --plugin-configs=NodeStatus run-plugins PciePluginUsing a JSON file:
node-scraper --plugin-configs=plugin_config.jsonHere is an example of a comprehensive plugin config that specifies analyzer args for each plugin:
{
"global_args": {},
"plugins": {
"BiosPlugin": {
"analysis_args": {
"exp_bios_version": "3.5"
}
},
"CmdlinePlugin": {
"analysis_args": {
"cmdline": "imgurl=test NODE=nodename selinux=0 serial console=ttyS1,115200 console=tty0",
"required_cmdline" : "selinux=0"
}
},
"DkmsPlugin": {
"analysis_args": {
"dkms_status": "amdgpu/6.11",
"dkms_version" : "dkms-3.1",
"regex_match" : true
}
},
"KernelPlugin": {
"analysis_args": {
"exp_kernel": "5.11-generic"
}
},
"OsPlugin": {
"analysis_args": {
"exp_os": "Ubuntu 22.04.2 LTS"
}
},
"PackagePlugin": {
"analysis_args": {
"exp_package_ver": {
"gcc": "11.4.0"
},
"regex_match": false
}
},
"RocmPlugin": {
"analysis_args": {
"exp_rocm": "6.5"
}
}
},
"result_collators": {},
"name": "plugin_config",
"desc": "My golden config"
}Post-action plugins run automatically after all primary plugins have completed, but only when one or more configurable conditions are met. They are defined in the same plugin config JSON as the primary plugins, under the post_action_plugins key.
Use cases:
{
"plugins": { ... },
"post_action_plugins": [
{
"plugin": "<PluginName>",
"plugin_args": { ... },
"conditions": [
{ "<field>": "<value>", ... },
{ "<field>": "<value>", ... }
]
}
]
}All fields are optional. A condition with no fields specified matches any result.
| Field | Type | Description |
|---|---|---|
| plugin | string | If set, only the result whose source matches this name is inspected. If omitted, all primary results are candidates. |
| status | string | The primary plugin's ExecutionStatus must be ≥ this value. Accepted values (in ascending order): OK, WARNING, ERROR, EXECUTION_FAILURE. |
| event_category | string | At least one event (from analysis or collection) must have this category. Normalised to uppercase with spaces/hyphens converted to underscores before comparison. |
| event_priority | string | At least one event's priority must be ≥ this value. Accepted values: INFO, WARNING, ERROR, CRITICAL. |
| event_description_contains | string | At least one event's description must contain this substring (case-sensitive). |
{
"name": "DmesgWithOsPostAction",
"desc": "Run DmesgPlugin; if any error-level event is found, run OsPlugin to capture OS state.",
"global_args": {},
"plugins": {
"DmesgPlugin": {
"collection": true,
"analysis": true
}
},
"result_collators": {},
"post_action_plugins": [
{
"plugin": "OsPlugin",
"plugin_args": {
"collection": true,
"analysis": true
},
"conditions": [
{
"plugin": "DmesgPlugin",
"event_priority": "ERROR"
}
]
}
]
}Save to a file and pass it with --plugin-configs:
node-scraper --plugin-configs=plugin_config_dmesg_os_post_action.jsonPost-action plugin results are included in the same result list as primary plugins — they appear in the console summary table, the nodescraper.csv output, and any result hooks.
Note: Post-action plugins run before connections are closed, so they have access to the same live connection managers as primary plugins. Post-action plugins cannot enqueue additional plugins into the primary queue.
This command can be used to generate a reference config that is populated with current system configurations. Plugins that use analyzer args (where applicable) will be populated with system data.
Run all registered in-band plugins (AllIbPlugins config):
node-scraper --plugin-configs=AllIbPlugins
Generate a reference config for specific plugins:
node-scraper --gen-reference-config run-plugins BiosPlugin OsPluginThis will generate the following config:
{
"global_args": {},
"plugins": {
"BiosPlugin": {
"analysis_args": {
"exp_bios_version": [
"M17"
],
"regex_match": false
}
},
"OsPlugin": {
"analysis_args": {
"exp_os": [
"8.10"
],
"exact_match": true
}
}
},
"result_collators": {}This config can later be used on a different platform for comparison, using the steps at #2:
node-scraper --plugin-configs=reference_config.json
An alternate way to generate a reference config is by using log files from a previous run. The example below uses log files from 'scraper_logs_/':
node-scraper gen-plugin-config --gen-reference-config-from-logs scraper_logs_<path>/ --output-path custom_output_dirThis will generate a reference config that includes plugins with logged results in 'scraper_log_' and save the new config to 'custom_output_dir/reference_config.json'.
| Back | FazBrowse Home | New Git URL |