diff --git a/docs/Collecting Metrics/Collectors/Applications/RUM Receiver.mdx b/docs/Collecting Metrics/Collectors/Applications/RUM Receiver.mdx index 06ff23a181..f63c5e4b52 100644 --- a/docs/Collecting Metrics/Collectors/Applications/RUM Receiver.mdx +++ b/docs/Collecting Metrics/Collectors/Applications/RUM Receiver.mdx @@ -67,6 +67,15 @@ DEM is experimental and opt-in. Build Netdata with `-DENABLE_PLUGIN_DEM=ON`, or Expose the listener through an HTTPS reverse proxy before sending traffic from remote websites. Set `public_url` and restrict `trusted_proxies` to that proxy. The default listener is local to this host. +#### Provision optional geography data + +Geography is optional. Provision an MMDB readable by the Netdata service account through Agent packaging, the topology IP intelligence downloader, or your own database update process. Some installations package these files with NetFlow; they are not present in every installation, and DEM does not require a running NetFlow plugin or download databases itself. + +Supported database types are `Netdata-Topology-GEO`, `GeoLite2-City`, `GeoLite2-Country`, `GeoIP2-City`, `GeoIP2-Country`, `DBIP-City-Lite` and `DBIP-Country-Lite`. ASN and other database types are rejected. A country database supplies countries without city coordinates; an accepted type does not guarantee coverage for an address. + +Publish updates by writing a separate complete file and atomically renaming it over the destination. Do not truncate or overwrite the active file: atomic replacement is required for consistent lookups. Check `collector.geoip` in `rum-sites` after provisioning. + + ### Configuration @@ -91,7 +100,13 @@ Options apply to the canonical receiver. | | rate_limit.per_ip_per_min | Maximum requests per client IP per minute. | 120 | no | | | rate_limit.per_site_per_sec | Maximum requests per site per second. | 500 | no | | **Collection** | update_every | Data collection interval, in seconds. | 10 | no | -| | geoip_db | MaxMind city database file. Leave empty to use the Agent IP intelligence database; unavailable data produces unknown locations. | | no | +| | [geoip_db](#option-collection-geoip-db) | Geographic MMDB file; an explicit path uses only that file. Leave empty to try the Agent cache and then stock IP intelligence database; unavailable geography does not stop collection. | | no | + + +##### geoip_db + +With an empty path, DEM tries `topology-ip-intel/topology-ip-geo.mmdb` under the Agent cache directory, then under its stock data directory. A rejected cache file does not hide an accepted stock file. An explicit path disables this automatic fallback. Sources are checked when the receiver starts and every 60 seconds, so a later provision or atomic replacement is adopted without restarting the receiver. + @@ -162,6 +177,21 @@ Metrics: ## Troubleshooting +### Known Errors + +#### Missing or unexpected geography + +**Cause** + +A site's capture policy can disable geography or limit it to country. The selected database may be unavailable, lack a matching address or field, or contain a record that cannot be decoded. + +**Fix** + +Check the site's `capture.geolocation` and `collector.geoip` in `rum-sites`. `loaded` means a supported declared database type and available reader, not full record validation, freshness, accuracy or coverage. `using_previous` means current candidates failed and a previous usable snapshot remains; confirmed source removal revokes it. `unavailable` does not stop browser collection. + +Use `reason` and the local Agent logs to identify source failures. Make a supported MMDB readable by the Netdata service account, then publish it by atomic rename. A snapshot with a detected mapped-memory fault is never reused. `lookup_errors` counts errors for the active source generation, not unmatched addresses or missing fields. Refresh and fault changes are published by the receiver worker; lookup counts update at the receiver collection interval. + + ### Other Problems #### No browser measurements diff --git a/docs/Collecting Metrics/Collectors/Hardware and Sensors/IPMI experimental Go collector.mdx b/docs/Collecting Metrics/Collectors/Hardware and Sensors/IPMI experimental Go collector.mdx index f9de80ff50..3f8ac2b3b2 100644 --- a/docs/Collecting Metrics/Collectors/Hardware and Sensors/IPMI experimental Go collector.mdx +++ b/docs/Collecting Metrics/Collectors/Hardware and Sensors/IPMI experimental Go collector.mdx @@ -29,7 +29,7 @@ Module: ipmi Monitor hardware sensor readings, sensor health and System Event Log entries through IPMI. -Queries the local Linux OpenIPMI device. The collector reads sensors and SEL metadata; it does not change BMC configuration or clear the event log. +Queries a local Linux OpenIPMI device or a remote BMC over IPMI 1.5 LAN or IPMI 2.0 LAN+ using the configured account. The collector reads sensors and SEL metadata; it does not change BMC configuration or clear the event log. This collector is only supported on the following platforms: @@ -37,7 +37,14 @@ This collector is only supported on the following platforms: This collector supports collecting metrics from multiple instances of this integration. -Local collection requires read/write access to the selected OpenIPMI device. Installation grants no setuid mode, capabilities or device permissions. +| Permission | Needed when | +|------------|-------------| +| Read/write access to the selected OpenIPMI device | Collecting locally with `driver: open`. | +| BMC account with User privilege and LAN access | Collecting remotely with `driver: lan` or `driver: lanplus`. Some BMC policies require a higher session privilege. | + +Installation grants no setuid mode, capabilities or device permissions. Remote collection requires no +elevated privileges on the Agent host. + ### Default Behavior @@ -47,7 +54,7 @@ No jobs start by default. Build the experimental plugin explicitly and configure #### Limits -This experimental plugin requires Linux amd64 or arm64 and supports local OpenIPMI only. Keep FreeIPMI for legacy direct drivers, OEM interpretation or custom FreeIPMI interpretation files. Unknown or unavailable sensor readings leave gaps. Default FreeIPMI policy treats non-critical threshold assertions as nominal. Do not run both IPMI plugins against the same host with the canonical ipmi job name because chart IDs overlap. The Go collector uses separate ipmi_go.* contexts and the ipmi-go-sensors Function. +This experimental plugin requires Linux amd64 or arm64. Keep FreeIPMI for legacy direct drivers, OEM interpretation or custom FreeIPMI interpretation files. Unknown or unavailable sensor readings leave gaps. Default FreeIPMI policy treats non-critical threshold assertions as nominal. Do not run both IPMI plugins against the same host with the canonical ipmi job name because chart IDs overlap. The Go collector uses separate ipmi_go.* contexts and the ipmi-go-sensors Function. #### Performance Impact @@ -64,24 +71,34 @@ Run `sudo ./netdata-installer.sh --enable-plugin-ipmi` from a source checkout. T #### Select one IPMI plugin -Before activating a Go IPMI job, set `freeipmi = no` in the `[plugins]` section of netdata.conf. Use a trusted configuration and explicit sudo invocation for local debugging when the device is root-only. For normal Agent execution, arrange device access according to the host security policy. +Before collecting the same host with the canonical ipmi job name, set `freeipmi = no` in the `[plugins]` section of netdata.conf. Use a trusted configuration and explicit sudo invocation for local debugging when the device is root-only. For normal Agent execution, arrange device access according to the host security policy. + +#### Enable remote BMC access + +For LAN/LAN+, enable IPMI over LAN on the BMC, configure an account with User privilege, and allow UDP traffic from the Agent to the BMC port (623 by default). Select a higher privilege only if the BMC access policy requires it. ### Configuration #### Options -Options apply to each local IPMI job. Remote LAN/LAN+ is not supported in this experimental build. +Options apply to each IPMI job. Configure one local device or remote BMC per job. | Group | Option | Description | Default | Required | |:------|:-----|:------------|:--------|:---------:| -| **Connection** | driver | Local Linux OpenIPMI transport. Only open is supported. | open | no | +| **Connection** | driver | IPMI transport: local Linux OpenIPMI, IPMI 1.5 LAN or IPMI 2.0 LAN+. | open | no | | | device | Local OpenIPMI device number. | 0 | no | +| | hostname | BMC hostname or IP address for LAN/LAN+, without a URL scheme or port. | | yes | +| | port | BMC UDP port for LAN/LAN+. | 623 | no | +| | username | BMC account for LAN/LAN+. Leave empty only for anonymous access. | | no | +| | password | BMC account password for LAN/LAN+. Leave empty only if the account has no password. | | no | +| | privilege_level | Session privilege requested from the BMC. User privilege is sufficient for standard sensor and SEL commands. | user | no | | **Collection** | update_every | Data collection frequency in seconds. | 5 | no | -| | timeout | Receive timeout for each IPMI command. Stopping a job can wait for an in-flight receive to finish. | 5s | no | +| | timeout | Timeout for session setup and for each IPMI command. | 5s | no | | | collect_sel | Collect the System Event Log entry count. | yes | no | +| **Virtual Node** | vnode | Associates this job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no | @@ -111,6 +128,33 @@ jobs: driver: open ``` +###### Remote BMC + +Monitor a remote BMC using IPMI 2.0. Use driver lan for BMCs that require IPMI 1.5. + +```yaml +jobs: + - name: server-bmc + driver: lanplus + hostname: bmc.example + username: monitor + password: '${env:IPMI_PASSWORD}' + +``` +###### Attach a remote BMC to a virtual node + +Define server-node in vnodes.conf first. Without vnode, metrics belong to the Agent host. + +```yaml +jobs: + - name: server-bmc + driver: lanplus + hostname: bmc.example + username: monitor + password: '${env:IPMI_PASSWORD}' + vnode: server-node + +``` ## Alerts diff --git a/docs/Collecting Metrics/Collectors/Hardware and Sensors/NVIDIA Jetson.mdx b/docs/Collecting Metrics/Collectors/Hardware and Sensors/NVIDIA Jetson.mdx index 2c163ea8f8..fa1262f178 100644 --- a/docs/Collecting Metrics/Collectors/Hardware and Sensors/NVIDIA Jetson.mdx +++ b/docs/Collecting Metrics/Collectors/Hardware and Sensors/NVIDIA Jetson.mdx @@ -6,7 +6,7 @@ toc_max_heading_level: "6" toc_collapsible: "true" learn_rel_path: "Collecting Metrics/Collectors/Hardware and Sensors" keywords: [nvidia jetson, jetson, nvidia, tegra, tegrastats, l4t, orin, xavier, thor, emc] -description: "Monitor GPU and memory controller utilization and clock frequencies on NVIDIA Jetson devices." +description: "Monitor GPU and memory controller utilization, clock frequencies, and power rail consumption on NVIDIA Jetson devices." message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE" sidebar_position: "250" learn_link: "https://learn.netdata.cloud/docs/collecting-metrics/collectors/hardware-and-sensors/nvidia-jetson" @@ -27,15 +27,16 @@ Module: jetson ## Overview -Monitor GPU and memory controller utilization and clock frequencies on NVIDIA Jetson devices. +Monitor GPU and memory controller utilization, clock frequencies, and power rail consumption on NVIDIA Jetson devices. You get: - GPU utilization. - GPU clock frequency, as one value or, on devices that report it, one value per graphics processing cluster (GPC). - External memory controller (EMC) bandwidth utilization and clock frequency. The CPU and GPU share this memory, so high EMC utilization can reveal memory-bandwidth-bound workloads. +- Current power consumption for each reported power rail, in watts. -CPU, memory, and swap usage come from Netdata's standard Linux system monitoring; this collector adds the Jetson GPU and EMC readings those charts do not cover. +CPU, memory, and swap usage come from Netdata's standard Linux system monitoring; this collector adds Jetson GPU, EMC, and power rail readings. Power readings help correlate workload activity with power consumption. The collector runs NVIDIA's `tegrastats --interval 1000` on the local host and reads the line of readings it prints every second. One `tegrastats` process runs for as long as the job runs, under the Netdata service account (usually `netdata`); the collector never elevates its privileges. @@ -63,6 +64,8 @@ If `tegrastats` is missing when Netdata starts, the job does not start; restart #### Limits - The readings depend on the Jetson module, the Jetson Linux release, and permissions. For example, `tegrastats` on Jetson Thor can report GPU clocks without GPU utilization. A reading that `tegrastats` does not report leaves a gap; it is never charted as zero. +- Rail names and coverage vary by device. Rails can overlap, so adding their values does not give total device power. On Thor, the total-system `VIN` rail may also appear in Netdata's sensors charts. +- Power charts use the current reading from supported current/average pairs. Averages and unfamiliar three-value forms are not collected. - If `tegrastats` stops printing readings, the charts show a gap after three seconds instead of repeating the last values. @@ -101,7 +104,7 @@ Check that the Netdata service account can run it: sudo -u netdata tegrastats --interval 1000 ``` -Each line should contain a `GR3D_FREQ` (GPU) or `EMC_FREQ` (memory controller) reading; stop the command with Ctrl+C. Readings missing from this output are also missing from Netdata. +Each line should contain a `GR3D_FREQ` (GPU), `EMC_FREQ` (memory controller), or power rail reading such as `VDD_GPU 10687mW/9815mW`; stop the command with Ctrl+C. Readings missing from this output are also missing from Netdata. If the command is not found, restore `tegrastats` with NVIDIA's [tegrastats deployment instructions](https://docs.nvidia.com/jetson/archives/r39.2/DeveloperGuide/AT/JetsonLinuxDevelopmentTools/TegrastatsUtility.html#re-deploying-tegrastats). Netdata searches its own `PATH`, which adds `/sbin`, `/usr/sbin`, `/usr/local/bin`, and `/usr/local/sbin` to the `PATH` it starts with. If `tegrastats` is in another directory, set `PATH` in the `[environment variables]` section of `netdata.conf` to a list that includes it, then restart Netdata. @@ -229,6 +232,23 @@ Metrics: | jetson.gpu_gpc_frequency | GPU Graphics Processing Cluster Frequency | frequency | MHz | +### Per power rail + +One power rail reported by `tegrastats` with a valid current/average milliwatt pair. + +Labels: + +| Label | Description | +|:-----------|:----------------| +| rail | NVIDIA's power rail name, such as VDD_GPU, VDD_CPU_CV, or VIN; a rail can supply several components. | + +Metrics: + +| Metric | Description | Dimensions | Unit | +|:------|:------------|:----------|:----| +| jetson.power_rail_power | Power Rail Consumption | power | Watts | + + ## Troubleshooting @@ -343,12 +363,12 @@ Netdata has not received a complete `tegrastats` line in the last three seconds, If the error persists, look for `tegrastats source unavailable` messages in the collector log (see Diagnostics) and run the check in Prerequisites. -#### tegrastats sample contains no supported GPU or EMC readings +#### tegrastats sample contains no supported GPU, EMC or power readings **Cause** -The latest `tegrastats` line has neither a valid `GR3D_FREQ` (GPU) nor `EMC_FREQ` (memory controller) reading. Which readings `tegrastats` prints depends on the Jetson module, the Jetson Linux release, and the account that runs it. +The latest `tegrastats` line has no valid `GR3D_FREQ` (GPU), `EMC_FREQ` (memory controller), or supported power rail reading. Which readings `tegrastats` prints depends on the Jetson module, the Jetson Linux release, and the account that runs it. **Fix** -Run the check in Prerequisites. If its output has no `GR3D_FREQ` or `EMC_FREQ` readings, this collector has nothing to chart on this device. +Run the check in Prerequisites. If its output has no supported GPU, memory controller, or power rail readings, this collector has nothing to chart on this device. diff --git a/docs/Developer and Contributor Corner/Build the Netdata Agent Yourself/How to build native DEB RPM packages locally for testing.mdx b/docs/Developer and Contributor Corner/Build the Netdata Agent Yourself/How to build native DEB RPM packages locally for testing.mdx index 1c1e283ae2..ce786d4cd5 100644 --- a/docs/Developer and Contributor Corner/Build the Netdata Agent Yourself/How to build native DEB RPM packages locally for testing.mdx +++ b/docs/Developer and Contributor Corner/Build the Netdata Agent Yourself/How to build native DEB RPM packages locally for testing.mdx @@ -41,9 +41,6 @@ packages for. For example, `centos7-v2` to build on CentOS 7, or `ubuntu20.04-v2 to build on Ubuntu 20.04. Note that we use Rocky Linux for builds on CentOS/RHEL 8 or newer. See [netdata/package-builders](https://hub.docker.com/r/netdata/package-builders/tags) for all available tags. -The `-v1` tags for RPM based distributions are still published, but they build from `netdata.spec.in` instead of -CMake and CPack. They are retained only as a fallback and no longer match what CI builds. - The value passed in the `VERSION` environment variable can be any version number accepted by the type of package being built. As a general rule, it needs to start with a digit, and must include a `.` somewhere. diff --git a/docs/Developer and Contributor Corner/Health command API tester.mdx b/docs/Developer and Contributor Corner/Health command API tester.mdx index 94621eb86b..50dca73e7e 100644 --- a/docs/Developer and Contributor Corner/Health command API tester.mdx +++ b/docs/Developer and Contributor Corner/Health command API tester.mdx @@ -10,7 +10,7 @@ slug: "/developer-and-contributor-corner/health-command-api-tester" # Health command API tester -The directory `tests/health_cmdapi` contains the test script `health-cmdapi-test.sh` for the [health command API](/docs/alerts-&-notifications/health-api-calls). +The directory `tests/health_mgmtapi` contains the test script `health-cmdapi-test.sh.in` for the [health command API](/docs/alerts-&-notifications/health-api-calls). The build no longer generates it; substitute `@varlibdir_POST@` by hand to run it. The script can be executed with options to prepare the system for the tests, run them and restore the system to its previous state. diff --git a/docs/Netdata Agent/Installation/Linux/Linux.mdx b/docs/Netdata Agent/Installation/Linux/Linux.mdx index 16b9a95735..2cd6a55528 100644 --- a/docs/Netdata Agent/Installation/Linux/Linux.mdx +++ b/docs/Netdata Agent/Installation/Linux/Linux.mdx @@ -127,7 +127,7 @@ The user running the script needs write and execute permissions in the temporary Before running the installation script, you can verify its integrity using the following command: ```bash -[ "44bdd008153b0696d007b3ecf2b566b9" = "$(curl -Ss https://get.netdata.cloud/kickstart.sh | md5sum | cut -d ' ' -f 1)" ] && echo "OK, VALID" || echo "FAILED, INVALID" +[ "d12eb806ced610078da979cccd59deae" = "$(curl -Ss https://get.netdata.cloud/kickstart.sh | md5sum | cut -d ' ' -f 1)" ] && echo "OK, VALID" || echo "FAILED, INVALID" ``` If the script is valid, this command will return `OK, VALID`. We recommend verifying script integrity before installation, especially in production environments. diff --git a/docs/Welcome to Netdata/Monitor Anything.mdx b/docs/Welcome to Netdata/Monitor Anything.mdx index d51c3a899b..a1a2429dc7 100644 --- a/docs/Welcome to Netdata/Monitor Anything.mdx +++ b/docs/Welcome to Netdata/Monitor Anything.mdx @@ -422,7 +422,7 @@ Need a dedicated integration? [Submit a feature request](https://github.com/netd | [Netatmo sensors](/docs/collecting-metrics/collectors/hardware-and-sensors/netatmo-sensors) | Keep an eye on Netatmo smart home device metrics for efficient home automation and energy management. | | [Nvidia Data Center GPU Manager (DCGM)](/docs/collecting-metrics/collectors/hardware-and-sensors/nvidia-data-center-gpu-manager-dcgm) | This collector gathers NVIDIA GPU telemetry from a `dcgm-exporter` endpoint. | | [Nvidia GPU](/docs/collecting-metrics/collectors/hardware-and-sensors/nvidia-gpu) | This collector monitors GPUs performance metrics using the [nvidia-smi](https://developer.nvidia.com/nvidia-system-management-interface) CLI tool. | -| [NVIDIA Jetson](/docs/collecting-metrics/collectors/hardware-and-sensors/nvidia-jetson) | Monitor GPU and memory controller utilization and clock frequencies on NVIDIA Jetson devices. | +| [NVIDIA Jetson](/docs/collecting-metrics/collectors/hardware-and-sensors/nvidia-jetson) | Monitor GPU and memory controller utilization, clock frequencies, and power rail consumption on NVIDIA Jetson devices. | | [Personal Weather Station](/docs/collecting-metrics/collectors/hardware-and-sensors/personal-weather-station) | Track personal weather station metrics for efficient weather monitoring and management. | | [Philips Hue](/docs/collecting-metrics/collectors/hardware-and-sensors/philips-hue) | Keep an eye on Philips Hue smart lighting metrics for efficient home automation and energy management. | | [Pimoroni Enviro+](/docs/collecting-metrics/collectors/hardware-and-sensors/pimoroni-enviro+) | Track Pimoroni Enviro+ air quality and environmental metrics for efficient environmental monitoring and analysis. | diff --git a/ingest/generated_map.yaml b/ingest/generated_map.yaml index 568e6160c1..c8adf79070 100644 --- a/ingest/generated_map.yaml +++ b/ingest/generated_map.yaml @@ -3738,8 +3738,8 @@ learn_rel_path: Collecting Metrics/Collectors/Hardware and Sensors keywords: '[''nvidia jetson'', ''jetson'', ''nvidia'', ''tegra'', ''tegrastats'', ''l4t'', ''orin'', ''xavier'', ''thor'', ''emc'']' - description: Monitor GPU and memory controller utilization and clock frequencies - on NVIDIA Jetson devices. + description: Monitor GPU and memory controller utilization, clock frequencies, and + power rail consumption on NVIDIA Jetson devices. meta_yaml: https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/jetson/metadata.yaml message: DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE diff --git a/ingest/generated_sidebar_order.json b/ingest/generated_sidebar_order.json index 10a9eef336..ffca2387bc 100644 --- a/ingest/generated_sidebar_order.json +++ b/ingest/generated_sidebar_order.json @@ -2,7 +2,7 @@ "schema_version": 1, "source": "netdata/docs/.map/map.yaml", "source_sha256": "5c613e8fa52bfee42b4c57d2ad18bf26ad3c1865ca5a5654eca0c66042d0c280", - "source_corpus_sha256": "d72c4b3e57a46a562a68650aba79490768b5f38ace7af85136329d0fbbc93363", + "source_corpus_sha256": "1fd618f3d43e51cc55df3c6204479305c89a5fb50d5fb53ce7f2e966fa71c57c", "order": [ { "parent_path": "Alerts & Notifications", @@ -1810,5 +1810,5 @@ "position": 190 } ], - "full_ingest_identity_sha256": "1bb44a6e40c5e715c2d71702fdf406be28d2f3c442dacfa403277fc5e176fc09" + "full_ingest_identity_sha256": "32faf2bc7348732a4d18711eadc37fcd1702eb8fa6c32250a095ebccbaf024f2" } diff --git a/ingest/generated_sidebar_order.json.sha256 b/ingest/generated_sidebar_order.json.sha256 index eca50ffc1a..bf1a259753 100644 --- a/ingest/generated_sidebar_order.json.sha256 +++ b/ingest/generated_sidebar_order.json.sha256 @@ -1 +1 @@ -4dec07c5b5b62bda0351157d58e951df880787bb51ed7d4a4c4a036d939fbaa8 generated_sidebar_order.json +1935f7283a5ae8b8ace105a70682e3694535051916b8e8eb9da5d4c38e805ab4 generated_sidebar_order.json