Showing posts with label FR511. Show all posts
Showing posts with label FR511. Show all posts

Tuesday, July 19, 2011

SNMP FAQ

Q. Why using a named pipe?
A. Because the healthcheck socket port has to be known by both the SNMP agent and WATCH, and this way this is transparent for the user (otherwise configureSnmp would have needed to be updated to ask (still) another port to the user, which was not a good solution).

Q. Why a double (named pipe / socket) mechanism to exchange information with the agent?
A. The named pipe / socket pair has been used because a full duplex named pipe (which would have seemed a simple solution) is not possible due to blocking issues in java.

Q. Does it perturb the SNMP agent classical queries?
A. Tests have shown that the additional 1 ms. timeout per query in the agent is not perceptible by the user.

Q. Why does the SNMP agent re-transmit the healthcheck socket port every (WATCH health cycle – 1) requests?
A. If WATCH ever crashes, and the SNMP agent would never re-transmit the port, the health check of the SNMP agent would never be able again.
(WATCH health cycle – 1) is used because it allows re-establishing the connection just below the health cycle allowance of WATCH.

Q. If WATCH cannot read (for any reason) from the named pipe every time the SNMP agent writes into it, what about the information stored in the pipe (garbage)?
A. When WATCH reads from the named pipe, it reads all the information available from it. It the SNMP agent has written multiple times, then this information won’t be usable, correct. But next time the SNMP agent will write into it, the port information will be clean.

A few points about SNMP monitoring

Two levels:
1) SWCTRL is now health checked like any other process: part of the SW-WATCH-HEALTH-CHECK condition.
2) The SNMP agent itself is checked by WATCH: SW-SNMP condition.

Concerning the SNMP agent, the nominal behavior is the following.

WATCH side:
1) Sends an empty datagram to the SNMP agent to wake it up
3) Reads healthcheck socket port from named pipe /tmp/swagent.watch.SHM (timeout of 1ms.)
4) Sends datagram on healthcheck socket
7) Reads response on healthcheck socket from SNMP agent (timeout of 1ms.)

SNMP side:
2) If first time or (WATCH health cycle – 1 reached), open healthcheck socket and send connection port to named pipe /tmp/swagent.watch.SHM
5) Reads datagram received on healthcheck socket (timeout of 1 ms.)
6) Sends back datagram on healthcheck socket

About WATCH health check

  • Processes mentioned in omni_health.SHM but not configured are not taken into account for the SW-WATCH-HEALTH-CHECK condition (e.g. IPMG if IUP not configured).
  • If a health checked process crashes, WATCH will not complain: no event (feature / bug?). However, the SW-WATCH-HEALTH-CHECK condition will then be raised.

About resmon

  • resmon is not started if the FR425 node is enabled.
  • The way resources are monitored is different on Solaris and Linux: files m_sunos5.c / m_linux.c.
  • resmon has no shared memory.
  • To enable traces in resmon, use non-public parameter RESMON_TRACE (e.g. RESMON_TRACE=0xff) in omni_conf_info file. Traces are put into $OMNI_HOME/Logs/Event…

To compute available memory on Linux, resmon extracts information from /proc/meminfo:

 memory in use = 100 – 100 * available / total
 available = free + shared + buffers + cached
 total = used + free

solaris resmon

Salut Laurent,

Je viens de me rappeler avoir envoyé ce mail a Jacques il y a quelque temps pour des détails sur resmon sous Solaris.
Je joins le fichier exemple (solaris_resmon.c) qui se compile facilement avec:

gcc -o solaris_resmon solaris_resmon.c -lkstat

Le concentre de l’algo de resmon sous Solaris y est.
NB: Je tiens a rappeler que l’algo ne sort pas de mon cerveau (malade) mais du code source de top ;)

Maintenant pour retrouver ces chiffres avec des outils standards:
  • La memoire totale disponible se voit par exemple avec:
prtconf | grep Memory
  • La memoire disponible à l’ instant t peut se voir par exemple avec:
vmstat 1 10

sous la colonne “free” a partir de la 2eme ligne.

jerome


Hello Jacques,

The resmon conditions are here to detect over-utilization of the computer resources.
So, if the concerned resource is below the specified (or default) threshold, the resmon condition as shown by DISPLAY-PLATFORM-STATUS is set to TRUE, and as soon as the threshold has been reached, the resmon condition becomes FALSE, which results in the CE becoming "not fully operational": this is what the PM event shows.
When memory is reported to exceed 80% utilization like in the customer case, in no way this means that new calls would be rejected by SW. This only means that the machine memory is heavily used.
Now to check the calculation, this is unfortunately a bit more complicated under Solaris than under Linux. Under Linux, the information is read from /proc/meminfo. Under Solaris, the information is read from the kernel via the native kstat library.

Globally the formula for the memory is :
memory in use = 100 - 100 * available / total

To have an idea of how the amount of available memory and total memory are computed on Solaris, please have a look at the attached sample file which is an extract from the original resmon.c. You can compile it under

Solaris 10, and it will show you these two values for your machine.

Now, to turn this off, as explained in the resmon man-page, the operator

can specify
RESMON_MEMORY_THRESHOLD=0
in his omni_conf_info file.

Hope this helps.

Rgds,
jerome