RackMon
Vendor guide

HPE iLO temperature monitoring over SNMP

Every ProLiant measures the air going into it. Getting that number out over SNMP is straightforward once you know two things — and both of them are things HPE does differently from everyone else.

The trap, up front: iLO gives you a code, not a name.

Dell tells you a probe is called “System Board Inlet Temp”. HPE tells you it is 11. The sensor you want is the one whose cpqHeTemperatureLocale equals 11 — ambient; everything else is CPU, memory or power supply, and will alarm for reasons that have nothing to do with your room. Pick the wrong index and you get a plausible-looking number that is quietly measuring the wrong thing.

1. Turn the SNMP agent on in iLO

On some firmware the iLO SNMP agent ships disabled, so the first symptom is not a wrong reading — it is silence. In the iLO web interface the settings live under the management or SNMP section; the exact path moves between iLO 4, 5 and 6, so look for SNMP Settings rather than following a fixed menu trail. You need:

The same settings can be driven over the Redfish API or hponcfg if you would rather script it across a fleet than click through each server.

2. Find the ambient sensor

A ProLiant exposes a lot of temperature sensors — intake, one per CPU, memory zones, power supplies, sometimes storage cages. Only the ambient one tells you about the room. A CPU at 70 °C under load is perfectly normal and says nothing about your cooling; an intake at 32 °C says every machine in that rack is in trouble.

Walk the locale column and read the codes:

# where is each sensor?
snmpwalk -v2c -c public 10.0.0.25 1.3.6.1.4.1.232.6.2.6.8.1.3

# typical ProLiant output
...8.1.3.1.1 = INTEGER: 11    <- ambient. this is the one.
...8.1.3.1.2 = INTEGER: 6     <- cpu
...8.1.3.1.3 = INTEGER: 6     <- cpu
...8.1.3.1.4 = INTEGER: 7     <- memory
...8.1.3.1.5 = INTEGER: 10    <- power supply

The full enumeration, straight from the CPQHLTH-MIB:

ValueLocationValueLocation
1other8storage
2unknown9removable media
3system10power supply
4system board11ambient — the intake
5I/O board12chassis
6CPU13bridge card
7memory

On Gen8 and later the ambient sensor is usually number 1 — the one the iLO web interface labels 01-Inlet Ambient — but the numbering moves between models, and reading the locale column costs you one command.

Second trap: the index has two parts.

The table is indexed by chassis and then sensor, so a real instance ends in .1.1, not .1. An snmpget on …8.1.4.1 returns No Such Instance and sends people hunting for a firmware problem that does not exist. On a single-chassis server the chassis number is always 1.

3. Read the temperature

ObjectOIDNotes
cpqHeTemperatureCelsius1.3.6.1.4.1.232.6.2.6.8.1.4Whole degrees. No scaling. -1 means it could not be read.
cpqHeTemperatureLocale1.3.6.1.4.1.232.6.2.6.8.1.3Where the sensor is. 11 = ambient.
cpqHeTemperatureCondition1.3.6.1.4.1.232.6.2.6.8.1.62 = ok. Discard anything else.
cpqHeTemperatureThreshold1.3.6.1.4.1.232.6.2.6.8.1.5HPE's own limit for that sensor. Not an alert level — see below.
cpqHeTemperatureThresholdType1.3.6.1.4.1.232.6.2.6.8.1.7What crossing it does: 9 = caution, 15 = critical.
# the ambient sensor, chassis 1, sensor 1
snmpget -v2c -c public 10.0.0.25 1.3.6.1.4.1.232.6.2.6.8.1.4.1.1
# INTEGER: 22   ->   22 °C, no maths required

On SNMPv3, which is what you should be using if the management network is shared:

snmpget -v3 -l authPriv -u monitor \
  -a SHA -A '<auth-pass>' -x AES -X '<priv-pass>' \
  10.0.0.25 1.3.6.1.4.1.232.6.2.6.8.1.4.1.1

Unlike Dell, there is no divide-by-ten here: whole degrees, exactly as returned. HPE does define a cpqHeTemperatureHwLocation string at …8.1.8, but it is optional and only populated on complex multi-enclosure hardware — on an ordinary ProLiant it comes back empty, which is precisely why the locale code matters.

4. Check the condition before you trust the number

A sensor that has failed still answers; it just answers with something meaningless. Two guards, both cheap:

There is also a single system-wide summary at 1.3.6.1.4.1.232.6.2.6.3.0 (cpqHeThermalTempStatus), on the same 2 = ok scale. It is a useful coarse health check, but it tells you nothing about the room — a server can report ok while sitting in air that is slowly cooking the rack next to it.

5. Set a threshold that means something

It is tempting to read cpqHeTemperatureThreshold and alert on that number. Don't. It is the point at which HPE itself acts, and what it does depends on the sensor's threshold type:

Threshold typeValueWhat crossing it does
blowout5Fans in that area spin up to pull the temperature back down.
caution9Condition becomes degraded; the server continues or shuts down depending on its configured degraded action.
critical15Condition becomes failed and the server shuts down.

Alerting at that line means being told at the moment the machine powers itself off. You want to know long before. Check the operating specification for your exact ProLiant generation, but as a working starting point for a mixed server room, alert on the ambient reading:

Intake temperatureStatusWhat it means
up to 27 °CNominalWithin the ASHRAE recommended envelope.
27–32 °CWarningInvestigate airflow and cooling now, while you have time.
32 °C and aboveCriticalCooling is likely failing. Act.

Whatever numbers you choose, add a hold window — require the reading to stay over the line for a few minutes before it pages anyone. A single odd sample at 3 a.m. is not a cooling failure, and an alert that cries wolf gets muted, which is worse than no alert at all.

Doing this across a fleet

One ProLiant is a snmpget in a cron job. Thirty of them across four racks is a different problem: you want every intake reading in one place, the hot one obvious without opening thirty graphs, and one alert for a cooling failure rather than thirty.

Rack thermal map with each server drawn in its U position and coloured by intake temperature
Thirty ProLiants, four racks, one screen — the hot U is obvious without reading a single graph.

That is what RackMon does: the iLO profile above ships as a template, so adding a ProLiant sets up the ambient, CPU and fan metrics with the OIDs already right — including the two-part index and the fan speed column that is not the fan condition. It starts on sensor 1 because that is where the ambient probe normally is; confirm it with the locale walk above and, if your model differs, change the OID on that one device. Self-hosted in one container, free for unlimited devices.

See it with live data   Install it

Other hardware

Dell, Cisco, Juniper, APC and the vendor-neutral standard are all in the SNMP temperature OID reference, each with its own scaling and unit traps. There are step-by-step guides for Dell iDRAC — where the trap is tenths of a degree rather than a location code — and for APC cards and probes. If you are still deciding what to measure and where to put the sensors, start with the rack temperature monitoring guide.

← All vendor OIDs   Monitoring guide