Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 31 additions & 1 deletion docs/Collecting Metrics/Collectors/Applications/RUM Receiver.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,15 @@ DEM is experimental and opt-in. Build Netdata with `-DENABLE_PLUGIN_DEM=ON`, or

Expose the listener through an HTTPS reverse proxy before sending traffic from remote websites. Set `public_url` and restrict `trusted_proxies` to that proxy. The default listener is local to this host.

#### Provision optional geography data

Geography is optional. Provision an MMDB readable by the Netdata service account through Agent packaging, the topology IP intelligence downloader, or your own database update process. Some installations package these files with NetFlow; they are not present in every installation, and DEM does not require a running NetFlow plugin or download databases itself.

Supported database types are `Netdata-Topology-GEO`, `GeoLite2-City`, `GeoLite2-Country`, `GeoIP2-City`, `GeoIP2-Country`, `DBIP-City-Lite` and `DBIP-Country-Lite`. ASN and other database types are rejected. A country database supplies countries without city coordinates; an accepted type does not guarantee coverage for an address.

Publish updates by writing a separate complete file and atomically renaming it over the destination. Do not truncate or overwrite the active file: atomic replacement is required for consistent lookups. Check `collector.geoip` in `rum-sites` after provisioning.



### Configuration

Expand All @@ -91,7 +100,13 @@ Options apply to the canonical receiver.
| | rate_limit.per_ip_per_min | Maximum requests per client IP per minute. | 120 | no |
| | rate_limit.per_site_per_sec | Maximum requests per site per second. | 500 | no |
| **Collection** | update_every | Data collection interval, in seconds. | 10 | no |
| | geoip_db | MaxMind city database file. Leave empty to use the Agent IP intelligence database; unavailable data produces unknown locations. | | no |
| | [geoip_db](#option-collection-geoip-db) | Geographic MMDB file; an explicit path uses only that file. Leave empty to try the Agent cache and then stock IP intelligence database; unavailable geography does not stop collection. | | no |

<a id="option-collection-geoip-db"></a>
##### geoip_db

With an empty path, DEM tries `topology-ip-intel/topology-ip-geo.mmdb` under the Agent cache directory, then under its stock data directory. A rejected cache file does not hide an accepted stock file. An explicit path disables this automatic fallback. Sources are checked when the receiver starts and every 60 seconds, so a later provision or atomic replacement is adopted without restarting the receiver.



</details>
Expand Down Expand Up @@ -162,6 +177,21 @@ Metrics:

## Troubleshooting

### Known Errors

#### Missing or unexpected geography

**Cause**

A site's capture policy can disable geography or limit it to country. The selected database may be unavailable, lack a matching address or field, or contain a record that cannot be decoded.

**Fix**

Check the site's `capture.geolocation` and `collector.geoip` in `rum-sites`. `loaded` means a supported declared database type and available reader, not full record validation, freshness, accuracy or coverage. `using_previous` means current candidates failed and a previous usable snapshot remains; confirmed source removal revokes it. `unavailable` does not stop browser collection.

Use `reason` and the local Agent logs to identify source failures. Make a supported MMDB readable by the Netdata service account, then publish it by atomic rename. A snapshot with a detected mapped-memory fault is never reused. `lookup_errors` counts errors for the active source generation, not unmatched addresses or missing fields. Refresh and fault changes are published by the receiver worker; lookup counts update at the receiver collection interval.


### Other Problems

#### No browser measurements
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -29,15 +29,22 @@ Module: ipmi

Monitor hardware sensor readings, sensor health and System Event Log entries through IPMI.

Queries the local Linux OpenIPMI device. The collector reads sensors and SEL metadata; it does not change BMC configuration or clear the event log.
Queries a local Linux OpenIPMI device or a remote BMC over IPMI 1.5 LAN or IPMI 2.0 LAN+ using the configured account. The collector reads sensors and SEL metadata; it does not change BMC configuration or clear the event log.

This collector is only supported on the following platforms:

- Linux

This collector supports collecting metrics from multiple instances of this integration.

Local collection requires read/write access to the selected OpenIPMI device. Installation grants no setuid mode, capabilities or device permissions.
| Permission | Needed when |
|------------|-------------|
| Read/write access to the selected OpenIPMI device | Collecting locally with `driver: open`. |
| BMC account with User privilege and LAN access | Collecting remotely with `driver: lan` or `driver: lanplus`. Some BMC policies require a higher session privilege. |

Installation grants no setuid mode, capabilities or device permissions. Remote collection requires no
elevated privileges on the Agent host.


### Default Behavior

Expand All @@ -47,7 +54,7 @@ No jobs start by default. Build the experimental plugin explicitly and configure

#### Limits

This experimental plugin requires Linux amd64 or arm64 and supports local OpenIPMI only. Keep FreeIPMI for legacy direct drivers, OEM interpretation or custom FreeIPMI interpretation files. Unknown or unavailable sensor readings leave gaps. Default FreeIPMI policy treats non-critical threshold assertions as nominal. Do not run both IPMI plugins against the same host with the canonical ipmi job name because chart IDs overlap. The Go collector uses separate ipmi_go.* contexts and the ipmi-go-sensors Function.
This experimental plugin requires Linux amd64 or arm64. Keep FreeIPMI for legacy direct drivers, OEM interpretation or custom FreeIPMI interpretation files. Unknown or unavailable sensor readings leave gaps. Default FreeIPMI policy treats non-critical threshold assertions as nominal. Do not run both IPMI plugins against the same host with the canonical ipmi job name because chart IDs overlap. The Go collector uses separate ipmi_go.* contexts and the ipmi-go-sensors Function.

#### Performance Impact

Expand All @@ -64,24 +71,34 @@ Run `sudo ./netdata-installer.sh --enable-plugin-ipmi` from a source checkout. T

#### Select one IPMI plugin

Before activating a Go IPMI job, set `freeipmi = no` in the `[plugins]` section of netdata.conf. Use a trusted configuration and explicit sudo invocation for local debugging when the device is root-only. For normal Agent execution, arrange device access according to the host security policy.
Before collecting the same host with the canonical ipmi job name, set `freeipmi = no` in the `[plugins]` section of netdata.conf. Use a trusted configuration and explicit sudo invocation for local debugging when the device is root-only. For normal Agent execution, arrange device access according to the host security policy.

#### Enable remote BMC access

For LAN/LAN+, enable IPMI over LAN on the BMC, configure an account with User privilege, and allow UDP traffic from the Agent to the BMC port (623 by default). Select a higher privilege only if the BMC access policy requires it.


### Configuration

#### Options

Options apply to each local IPMI job. Remote LAN/LAN+ is not supported in this experimental build.
Options apply to each IPMI job. Configure one local device or remote BMC per job.



| Group | Option | Description | Default | Required |
|:------|:-----|:------------|:--------|:---------:|
| **Connection** | driver | Local Linux OpenIPMI transport. Only open is supported. | open | no |
| **Connection** | driver | IPMI transport: local Linux OpenIPMI, IPMI 1.5 LAN or IPMI 2.0 LAN+. | open | no |
| | device | Local OpenIPMI device number. | 0 | no |
| | hostname | BMC hostname or IP address for LAN/LAN+, without a URL scheme or port. | | yes |
| | port | BMC UDP port for LAN/LAN+. | 623 | no |
| | username | BMC account for LAN/LAN+. Leave empty only for anonymous access. | | no |
| | password | BMC account password for LAN/LAN+. Leave empty only if the account has no password. | | no |
| | privilege_level | Session privilege requested from the BMC. User privilege is sufficient for standard sensor and SEL commands. | user | no |
| **Collection** | update_every | Data collection frequency in seconds. | 5 | no |
| | timeout | Receive timeout for each IPMI command. Stopping a job can wait for an in-flight receive to finish. | 5s | no |
| | timeout | Timeout for session setup and for each IPMI command. | 5s | no |
| | collect_sel | Collect the System Event Log entry count. | yes | no |
| **Virtual Node** | vnode | Associates this job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |



Expand Down Expand Up @@ -111,6 +128,33 @@ jobs:
driver: open

```
###### Remote BMC

Monitor a remote BMC using IPMI 2.0. Use driver lan for BMCs that require IPMI 1.5.

```yaml
jobs:
- name: server-bmc
driver: lanplus
hostname: bmc.example
username: monitor
password: '${env:IPMI_PASSWORD}'

```
###### Attach a remote BMC to a virtual node

Define server-node in vnodes.conf first. Without vnode, metrics belong to the Agent host.

```yaml
jobs:
- name: server-bmc
driver: lanplus
hostname: bmc.example
username: monitor
password: '${env:IPMI_PASSWORD}'
vnode: server-node

```


## Alerts
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ toc_max_heading_level: "6"
toc_collapsible: "true"
learn_rel_path: "Collecting Metrics/Collectors/Hardware and Sensors"
keywords: [nvidia jetson, jetson, nvidia, tegra, tegrastats, l4t, orin, xavier, thor, emc]
description: "Monitor GPU and memory controller utilization and clock frequencies on NVIDIA Jetson devices."
description: "Monitor GPU and memory controller utilization, clock frequencies, and power rail consumption on NVIDIA Jetson devices."
message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
sidebar_position: "250"
learn_link: "https://learn.netdata.cloud/docs/collecting-metrics/collectors/hardware-and-sensors/nvidia-jetson"
Expand All @@ -27,15 +27,16 @@ Module: jetson

## Overview

Monitor GPU and memory controller utilization and clock frequencies on NVIDIA Jetson devices.
Monitor GPU and memory controller utilization, clock frequencies, and power rail consumption on NVIDIA Jetson devices.

You get:

- GPU utilization.
- GPU clock frequency, as one value or, on devices that report it, one value per graphics processing cluster (GPC).
- External memory controller (EMC) bandwidth utilization and clock frequency. The CPU and GPU share this memory, so high EMC utilization can reveal memory-bandwidth-bound workloads.
- Current power consumption for each reported power rail, in watts.

CPU, memory, and swap usage come from Netdata's standard Linux system monitoring; this collector adds the Jetson GPU and EMC readings those charts do not cover.
CPU, memory, and swap usage come from Netdata's standard Linux system monitoring; this collector adds Jetson GPU, EMC, and power rail readings. Power readings help correlate workload activity with power consumption.


The collector runs NVIDIA's `tegrastats --interval 1000` on the local host and reads the line of readings it prints every second. One `tegrastats` process runs for as long as the job runs, under the Netdata service account (usually `netdata`); the collector never elevates its privileges.
Expand Down Expand Up @@ -63,6 +64,8 @@ If `tegrastats` is missing when Netdata starts, the job does not start; restart
#### Limits

- The readings depend on the Jetson module, the Jetson Linux release, and permissions. For example, `tegrastats` on Jetson Thor can report GPU clocks without GPU utilization. A reading that `tegrastats` does not report leaves a gap; it is never charted as zero.
- Rail names and coverage vary by device. Rails can overlap, so adding their values does not give total device power. On Thor, the total-system `VIN` rail may also appear in Netdata's sensors charts.
- Power charts use the current reading from supported current/average pairs. Averages and unfamiliar three-value forms are not collected.
- If `tegrastats` stops printing readings, the charts show a gap after three seconds instead of repeating the last values.


Expand Down Expand Up @@ -101,7 +104,7 @@ Check that the Netdata service account can run it:
sudo -u netdata tegrastats --interval 1000
```

Each line should contain a `GR3D_FREQ` (GPU) or `EMC_FREQ` (memory controller) reading; stop the command with Ctrl+C. Readings missing from this output are also missing from Netdata.
Each line should contain a `GR3D_FREQ` (GPU), `EMC_FREQ` (memory controller), or power rail reading such as `VDD_GPU 10687mW/9815mW`; stop the command with Ctrl+C. Readings missing from this output are also missing from Netdata.

If the command is not found, restore `tegrastats` with NVIDIA's [tegrastats deployment instructions](https://docs.nvidia.com/jetson/archives/r39.2/DeveloperGuide/AT/JetsonLinuxDevelopmentTools/TegrastatsUtility.html#re-deploying-tegrastats). Netdata searches its own `PATH`, which adds `/sbin`, `/usr/sbin`, `/usr/local/bin`, and `/usr/local/sbin` to the `PATH` it starts with. If `tegrastats` is in another directory, set `PATH` in the `[environment variables]` section of `netdata.conf` to a list that includes it, then restart Netdata.

Expand Down Expand Up @@ -229,6 +232,23 @@ Metrics:
| jetson.gpu_gpc_frequency | GPU Graphics Processing Cluster Frequency | frequency | MHz |


### Per power rail

One power rail reported by `tegrastats` with a valid current/average milliwatt pair.

Labels:

| Label | Description |
|:-----------|:----------------|
| rail | NVIDIA's power rail name, such as VDD_GPU, VDD_CPU_CV, or VIN; a rail can supply several components. |

Metrics:

| Metric | Description | Dimensions | Unit |
|:------|:------------|:----------|:----|
| jetson.power_rail_power | Power Rail Consumption | power | Watts |



## Troubleshooting

Expand Down Expand Up @@ -343,12 +363,12 @@ Netdata has not received a complete `tegrastats` line in the last three seconds,

If the error persists, look for `tegrastats source unavailable` messages in the collector log (see Diagnostics) and run the check in Prerequisites.

#### tegrastats sample contains no supported GPU or EMC readings
#### tegrastats sample contains no supported GPU, EMC or power readings

**Cause**

The latest `tegrastats` line has neither a valid `GR3D_FREQ` (GPU) nor `EMC_FREQ` (memory controller) reading. Which readings `tegrastats` prints depends on the Jetson module, the Jetson Linux release, and the account that runs it.
The latest `tegrastats` line has no valid `GR3D_FREQ` (GPU), `EMC_FREQ` (memory controller), or supported power rail reading. Which readings `tegrastats` prints depends on the Jetson module, the Jetson Linux release, and the account that runs it.

**Fix**

Run the check in Prerequisites. If its output has no `GR3D_FREQ` or `EMC_FREQ` readings, this collector has nothing to chart on this device.
Run the check in Prerequisites. If its output has no supported GPU, memory controller, or power rail readings, this collector has nothing to chart on this device.
Original file line number Diff line number Diff line change
Expand Up @@ -41,9 +41,6 @@ packages for. For example, `centos7-v2` to build on CentOS 7, or `ubuntu20.04-v2
to build on Ubuntu 20.04. Note that we use Rocky Linux for builds on CentOS/RHEL 8 or newer. See
[netdata/package-builders](https://hub.docker.com/r/netdata/package-builders/tags) for all available tags.

The `-v1` tags for RPM based distributions are still published, but they build from `netdata.spec.in` instead of
CMake and CPack. They are retained only as a fallback and no longer match what CI builds.

The value passed in the `VERSION` environment variable can be any version number accepted by the type of package
being built. As a general rule, it needs to start with a digit, and must include a `.` somewhere.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ slug: "/developer-and-contributor-corner/health-command-api-tester"

# Health command API tester

The directory `tests/health_cmdapi` contains the test script `health-cmdapi-test.sh` for the [health command API](/docs/alerts-&-notifications/health-api-calls).
The directory `tests/health_mgmtapi` contains the test script `health-cmdapi-test.sh.in` for the [health command API](/docs/alerts-&-notifications/health-api-calls). The build no longer generates it; substitute `@varlibdir_POST@` by hand to run it.

The script can be executed with options to prepare the system for the tests, run them and restore the system to its previous state.

Expand Down
2 changes: 1 addition & 1 deletion docs/Netdata Agent/Installation/Linux/Linux.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -127,7 +127,7 @@ The user running the script needs write and execute permissions in the temporary
Before running the installation script, you can verify its integrity using the following command:

```bash
[ "44bdd008153b0696d007b3ecf2b566b9" = "$(curl -Ss https://get.netdata.cloud/kickstart.sh | md5sum | cut -d ' ' -f 1)" ] && echo "OK, VALID" || echo "FAILED, INVALID"
[ "d12eb806ced610078da979cccd59deae" = "$(curl -Ss https://get.netdata.cloud/kickstart.sh | md5sum | cut -d ' ' -f 1)" ] && echo "OK, VALID" || echo "FAILED, INVALID"
```

If the script is valid, this command will return `OK, VALID`. We recommend verifying script integrity before installation, especially in production environments.
Expand Down
2 changes: 1 addition & 1 deletion docs/Welcome to Netdata/Monitor Anything.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -422,7 +422,7 @@ Need a dedicated integration? [Submit a feature request](https://github.com/netd
| [Netatmo sensors](/docs/collecting-metrics/collectors/hardware-and-sensors/netatmo-sensors) | Keep an eye on Netatmo smart home device metrics for efficient home automation and energy management. |
| [Nvidia Data Center GPU Manager (DCGM)](/docs/collecting-metrics/collectors/hardware-and-sensors/nvidia-data-center-gpu-manager-dcgm) | This collector gathers NVIDIA GPU telemetry from a `dcgm-exporter` endpoint. |
| [Nvidia GPU](/docs/collecting-metrics/collectors/hardware-and-sensors/nvidia-gpu) | This collector monitors GPUs performance metrics using the [nvidia-smi](https://developer.nvidia.com/nvidia-system-management-interface) CLI tool. |
| [NVIDIA Jetson](/docs/collecting-metrics/collectors/hardware-and-sensors/nvidia-jetson) | Monitor GPU and memory controller utilization and clock frequencies on NVIDIA Jetson devices. |
| [NVIDIA Jetson](/docs/collecting-metrics/collectors/hardware-and-sensors/nvidia-jetson) | Monitor GPU and memory controller utilization, clock frequencies, and power rail consumption on NVIDIA Jetson devices. |
| [Personal Weather Station](/docs/collecting-metrics/collectors/hardware-and-sensors/personal-weather-station) | Track personal weather station metrics for efficient weather monitoring and management. |
| [Philips Hue](/docs/collecting-metrics/collectors/hardware-and-sensors/philips-hue) | Keep an eye on Philips Hue smart lighting metrics for efficient home automation and energy management. |
| [Pimoroni Enviro+](/docs/collecting-metrics/collectors/hardware-and-sensors/pimoroni-enviro+) | Track Pimoroni Enviro+ air quality and environmental metrics for efficient environmental monitoring and analysis. |
Expand Down
4 changes: 2 additions & 2 deletions ingest/generated_map.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -3738,8 +3738,8 @@
learn_rel_path: Collecting Metrics/Collectors/Hardware and Sensors
keywords: '[''nvidia jetson'', ''jetson'', ''nvidia'', ''tegra'', ''tegrastats'',
''l4t'', ''orin'', ''xavier'', ''thor'', ''emc'']'
description: Monitor GPU and memory controller utilization and clock frequencies
on NVIDIA Jetson devices.
description: Monitor GPU and memory controller utilization, clock frequencies, and
power rail consumption on NVIDIA Jetson devices.
meta_yaml: https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/jetson/metadata.yaml
message: DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml
FILE
Expand Down
Loading
Loading