Dell iDRAC temperature monitoring over SNMP
Every PowerEdge already measures its own inlet air. Getting that number out over SNMP takes about ten minutes — and there is one trap that catches almost everybody the first time.
A returned 275 means 27.5 °C. Divide by 10. This bites in both directions — a naive threshold of 35 will never fire, because the reading is always in the hundreds, and a dashboard that shows “275 °C” gets ignored as obviously broken.
1. Turn SNMP on in iDRAC
SNMP is not enabled out of the box on every firmware. In the iDRAC web interface it lives under the network or services settings — the exact path moves between iDRAC 7, 8 and 9, so look for SNMP Agent rather than following a fixed menu trail. You need three things:
- The SNMP agent enabled.
- An SNMPv3 user, or a v2c community string if you are on an isolated management VLAN and accept that it travels in clear text.
- Your monitoring host able to reach the iDRAC on UDP 161 — this is the step that most often turns out to be a firewall or VLAN problem rather than an iDRAC one.
The same settings are reachable over racadm if you would rather script it across a fleet than click through each card.
2. Find out which probe is the inlet
A PowerEdge exposes several temperature probes: inlet, exhaust, one per CPU, sometimes board and memory zones. Only the inlet tells you about the room. A CPU at 68 °C under load is fine and says nothing about your cooling; an inlet at 32 °C says every machine in that rack is in trouble.
Walk the names first, then match the index:
# the probe names
snmpwalk -v2c -c public 10.0.0.20 1.3.6.1.4.1.674.10892.5.4.700.20.1.8
# typical output
...20.1.8.1.1 = STRING: "System Board Inlet Temp"
...20.1.8.1.2 = STRING: "System Board Exhaust Temp"
...20.1.8.1.3 = STRING: "CPU1 Temp"
...20.1.8.1.4 = STRING: "CPU2 Temp"
Index 1 is normally the inlet, but do not assume it — it varies by model and configuration, and reading the names costs you one command.
3. Read the temperature
| Object | OID | Notes |
|---|---|---|
temperatureProbeReading | 1.3.6.1.4.1.674.10892.5.4.700.20.1.6 | Tenths of a degree. Divide by 10. |
temperatureProbeLocationName | 1.3.6.1.4.1.674.10892.5.4.700.20.1.8 | Which probe each index is. |
temperatureProbeStatus | 1.3.6.1.4.1.674.10892.5.4.700.20.1.5 | 3 = ok. Discard anything else. |
# the inlet, assuming index 1
snmpget -v2c -c public 10.0.0.20 1.3.6.1.4.1.674.10892.5.4.700.20.1.6.1.1
# INTEGER: 275 -> 27.5 °C
On SNMPv3, which is what you should be using if the management network is shared:
snmpget -v3 -l authPriv -u monitor \
-a SHA -A '<auth-pass>' -x AES -X '<priv-pass>' \
10.0.0.20 1.3.6.1.4.1.674.10892.5.4.700.20.1.6.1.1
4. Check the status before you trust the number
A probe that has failed or been removed still answers — it just answers with something meaningless. temperatureProbeStatus returns 3 when the reading is good. Poll it alongside the value and discard anything else, or you will eventually alert on a dead sensor reporting 0 °C and spend an hour looking for a cooling failure that never happened.
5. Set a threshold that means something
Alert on the inlet. Check the operating specification for your exact PowerEdge model, since it varies by generation and configuration — but as a working starting point for a mixed server room:
| Inlet temperature | Status | What it means |
|---|---|---|
| up to 27 °C | Nominal | Within the ASHRAE recommended envelope. |
| 27–32 °C | Warning | Investigate airflow and cooling now, while you have time. |
| 32 °C and above | Critical | Cooling is likely failing. Act. |
Dell also publishes its own per-probe thresholds over SNMP if you would rather use the vendor's numbers than pick your own. Whichever you choose, add a hold window — require the reading to stay over the line for a few minutes before it pages anyone. A single odd sample at 3 a.m. is not a cooling failure, and an alert that cries wolf gets muted, which is worse than no alert at all.
Doing this across a fleet
One server is a snmpget in a cron job. Forty servers across four racks is a different problem: you want the readings in one place, the hot one obvious without reading forty graphs, and an alert that fires once for a cooling failure rather than forty times.
That is what RackMon does: the iDRAC profile above is built in, including the divide-by-10, so you point it at a server and it finds the inlet probe itself. Self-hosted in one container, free for unlimited devices.
See it with live data Install it
Other hardware
Cisco, Juniper, APC and the vendor-neutral standard are all in the SNMP temperature OID reference, each with its own scaling and unit traps. There are step-by-step guides for HPE iLO — where the trap is a location code instead of a scaling factor — and for APC cards and probes too. If you are still deciding what to measure and where to put the sensors, start with the rack temperature monitoring guide.