1
1

00:00:00,240  -->  00:00:01,370
<v ->Sensors.</v>
2

2

00:00:01,370  -->  00:00:02,290
In this lesson,
3

3

00:00:02,290  -->  00:00:03,550
we're going to talk about sensors
4

4

00:00:03,550  -->  00:00:06,020
that help us monitor the performance of our network devices,
5

5

00:00:06,020  -->  00:00:09,080
those devices like routers, switches, and firewalls.
6

6

00:00:09,080  -->  00:00:10,810
Now, these sensors can be used to monitor
7

7

00:00:10,810  -->  00:00:14,110
the device's temperature, its CPU usage, and its memory,
8

8

00:00:14,110  -->  00:00:16,210
and these things can be key indicators
9

9

00:00:16,210  -->  00:00:18,080
of whether a device is operating properly
10

10

00:00:18,080  -->  00:00:21,180
or is about to suffer a catastrophic failure.
11

11

00:00:21,180  -->  00:00:23,290
Our first sensor measurement we need to talk about
12

12

00:00:23,290  -->  00:00:25,050
is the temperature of the device.
13

13

00:00:25,050  -->  00:00:27,130
Now, most network devices like your routers,
14

14

00:00:27,130  -->  00:00:28,280
switches, and firewalls
15

15

00:00:28,280  -->  00:00:30,350
have the ability to report on the temperature
16

16

00:00:30,350  -->  00:00:31,970
within their chasses.
17

17

00:00:31,970  -->  00:00:32,942
Now, depending on the model,
18

18

00:00:32,942  -->  00:00:35,570
there may be only one or two temperature readings
19

19

00:00:35,570  -->  00:00:37,900
or on some larger enterprise devices,
20

20

00:00:37,900  -->  00:00:39,280
you may have a temperature reading
21

21

00:00:39,280  -->  00:00:42,430
on each and every controller, processor, interface card,
22

22

00:00:42,430  -->  00:00:44,630
and thing like that inside the system.
23

23

00:00:44,630  -->  00:00:45,880
Now, the temperature centers
24

24

00:00:45,880  -->  00:00:47,530
can be used to measure the air temperature
25

25

00:00:47,530  -->  00:00:49,070
inside the intake outlet
26

26

00:00:49,070  -->  00:00:52,050
and the air temperature at the exhaust outlet at a minimum.
27

27

00:00:52,050  -->  00:00:53,360
Now, for each of these sensors,
28

28

00:00:53,360  -->  00:00:56,370
you can set up minor and major temperature thresholds.
29

29

00:00:56,370  -->  00:00:59,240
A minor temperature threshold is used to set off an alarm
30

30

00:00:59,240  -->  00:01:01,090
when a rising temperature is detected,
31

31

00:01:01,090  -->  00:01:03,630
but it hasn't reached dangerous levels yet.
32

32

00:01:03,630  -->  00:01:04,490
When this occurs,
33

33

00:01:04,490  -->  00:01:06,180
a system message is displayed,
34

34

00:01:06,180  -->  00:01:08,160
an SNMP notification is sent,
35

35

00:01:08,160  -->  00:01:10,730
and an environmental alarm can be sounded.
36

36

00:01:10,730  -->  00:01:12,940
Now, when you have a major temperature threshold,
37

37

00:01:12,940  -->  00:01:15,040
this is going to be used to set off an alarm
38

38

00:01:15,040  -->  00:01:17,510
when the temperature reaches dangerous conditions.
39

39

00:01:17,510  -->  00:01:20,790
At this level, we want to still display those system messages,
40

40

00:01:20,790  -->  00:01:22,880
get that SNMP notification,
41

41

00:01:22,880  -->  00:01:24,900
and have the environmental alarm sounded.
42

42

00:01:24,900  -->  00:01:26,400
But in addition to that,
43

43

00:01:26,400  -->  00:01:28,400
the device can actually start to load shed
44

44

00:01:28,400  -->  00:01:30,900
by turning off different functions to reduce the temperature
45

45

00:01:30,900  -->  00:01:33,300
being generated by the device's processor.
46

46

00:01:33,300  -->  00:01:35,020
For example, let's say you have a router
47

47

00:01:35,020  -->  00:01:37,020
with multiple processing cards in it.
48

48

00:01:37,020  -->  00:01:39,750
That device may shut down one of those processing cards
49

49

00:01:39,750  -->  00:01:42,010
to prevent the entire system from overheating.
50

50

00:01:42,010  -->  00:01:43,880
That's what I mean by load shedding.
51

51

00:01:43,880  -->  00:01:45,950
Now, when a device runs at excessive temperatures
52

52

00:01:45,950  -->  00:01:46,890
for too long,
53

53

00:01:46,890  -->  00:01:48,867
the performance will decrease on that device
54

54

00:01:48,867  -->  00:01:51,940
and the lifespan will decline on that device as well.
55

55

00:01:51,940  -->  00:01:54,010
Over time, that device can even suffer
56

56

00:01:54,010  -->  00:01:56,590
a catastrophic failure from overheating.
57

57

00:01:56,590  -->  00:01:58,740
Our second sensor measurement we need to talk about
58

58

00:01:58,740  -->  00:02:02,130
is CPU usage or utilization on the device.
59

59

00:02:02,130  -->  00:02:03,050
At their core,
60

60

00:02:03,050  -->  00:02:04,570
routers, switches, and firewalls
61

61

00:02:04,570  -->  00:02:06,420
are just specialized computers.
62

62

00:02:06,420  -->  00:02:08,760
When these devices are running under normal conditions,
63

63

00:02:08,760  -->  00:02:11,100
their CPU or central processing unit
64

64

00:02:11,100  -->  00:02:12,970
should have minimal utilization
65

65

00:02:12,970  -->  00:02:15,690
somewhere in the range of 5 to 40%.
66

66

00:02:15,690  -->  00:02:18,230
But if the devices begin to become extremely busy
67

67

00:02:18,230  -->  00:02:20,720
or receive too many packets from its neighboring devices,
68

68

00:02:20,720  -->  00:02:23,510
the CPU utilization can become over-utilized
69

69

00:02:23,510  -->  00:02:25,290
and the percentage will increase.
70

70

00:02:25,290  -->  00:02:27,820
Now, if the CPU utilization gets too high,
71

71

00:02:27,820  -->  00:02:30,850
the device could become unable to process any more requests
72

72

00:02:30,850  -->  00:02:32,360
and it'll start to drop packets
73

73

00:02:32,360  -->  00:02:34,870
or the entire connection could fail.
74

74

00:02:34,870  -->  00:02:37,880
Usually, when you see a high processor utilization rate,
75

75

00:02:37,880  -->  00:02:40,280
this is an indication of a misconfigured network
76

76

00:02:40,280  -->  00:02:42,010
or a network under attack.
77

77

00:02:42,010  -->  00:02:43,500
If the network is misconfigured,
78

78

00:02:43,500  -->  00:02:44,910
for example, let's say you have a switch
79

79

00:02:44,910  -->  00:02:45,980
that's misconfigured,
80

80

00:02:45,980  -->  00:02:48,450
you can end up having a broadcast storm that occurs,
81

81

00:02:48,450  -->  00:02:49,283
and that's going to create
82

82

00:02:49,283  -->  00:02:51,220
an excessive amount of broadcast traffic
83

83

00:02:51,220  -->  00:02:53,990
that'll cause the switch's CPU to become utilized
84

84

00:02:53,990  -->  00:02:56,460
as it tries to process all those requests.
85

85

00:02:56,460  -->  00:02:58,320
Similarly, if you have a lot of complex
86

86

00:02:58,320  -->  00:03:00,070
and intricate ACLs on your router,
87

87

00:03:00,070  -->  00:03:02,590
and then people started sending a lot of inbound traffic,
88

88

00:03:02,590  -->  00:03:05,620
that router has to go through all of those ACLs each time
89

89

00:03:05,620  -->  00:03:08,480
for that traffic and that can make it become unresponsive
90

90

00:03:08,480  -->  00:03:10,550
due to high CPU usage.
91

91

00:03:10,550  -->  00:03:11,650
As an administrator,
92

92

00:03:11,650  -->  00:03:13,550
you need to monitor the CPU utilization
93

93

00:03:13,550  -->  00:03:14,770
in your network devices
94

94

00:03:14,770  -->  00:03:16,650
to determine if they're operating properly,
95

95

00:03:16,650  -->  00:03:17,640
if they're misconfigured,
96

96

00:03:17,640  -->  00:03:19,530
or if they're under attack.
97

97

00:03:19,530  -->  00:03:21,160
The third sensor measurement we use
98

98

00:03:21,160  -->  00:03:23,550
is memory utilization for the device.
99

99

00:03:23,550  -->  00:03:25,550
Similar to high CPU utilization,
100

100

00:03:25,550  -->  00:03:27,630
high memory utilization can be indicative
101

101

00:03:27,630  -->  00:03:29,600
of a larger problem in your network.
102

102

00:03:29,600  -->  00:03:31,580
If your devices begin to use too much memory,
103

103

00:03:31,580  -->  00:03:34,200
this can lead to system hangs, processor crashes,
104

104

00:03:34,200  -->  00:03:36,230
and other undesirable behavior.
105

105

00:03:36,230  -->  00:03:37,640
To help protect against this,
106

106

00:03:37,640  -->  00:03:39,180
you should have minor, severe,
107

107

00:03:39,180  -->  00:03:41,040
and critical memory threshold warnings
108

108

00:03:41,040  -->  00:03:42,330
set up in your devices
109

109

00:03:42,330  -->  00:03:44,860
and reporting back to your centralized monitoring dashboard
110

110

00:03:44,860  -->  00:03:46,680
using SNMP.
111

111

00:03:46,680  -->  00:03:47,710
As a baseline,
112

112

00:03:47,710  -->  00:03:49,010
your never devices should operate
113

113

00:03:49,010  -->  00:03:51,010
at around 40% memory utilization
114

114

00:03:51,010  -->  00:03:52,940
under normal working conditions.
115

115

00:03:52,940  -->  00:03:54,120
During busier times,
116

116

00:03:54,120  -->  00:03:56,670
you may see this rise up to 60 to 70%,
117

117

00:03:56,670  -->  00:04:00,050
and during peak times it may be up to 80%,
118

118

00:04:00,050  -->  00:04:02,440
but if you're constantly seeing memory utilization
119

119

00:04:02,440  -->  00:04:03,910
above 80%,
120

120

00:04:03,910  -->  00:04:05,490
you may need to install a larger
121

121

00:04:05,490  -->  00:04:07,800
or more powerful device for your network,
122

122

00:04:07,800  -->  00:04:09,690
or you could be under an attack
123

123

00:04:09,690  -->  00:04:11,050
for an excessive amount of time
124

124

00:04:11,050  -->  00:04:12,950
that's causing excessive loading.
125

125

00:04:12,950  -->  00:04:15,340
As you begin to operate your networks in the real world,
126

126

00:04:15,340  -->  00:04:17,330
you're going to begin to see what normal looks like
127

127

00:04:17,330  -->  00:04:19,000
for your particular network.
128

128

00:04:19,000  -->  00:04:20,700
As you see temperatures rising
129

129

00:04:20,700  -->  00:04:23,240
or CPU and memory utilizations increase,
130

130

00:04:23,240  -->  00:04:25,680
this can trigger alarms in the network configuration
131

131

00:04:25,680  -->  00:04:28,870
or a network performance issue is happening right now.
132

132

00:04:28,870  -->  00:04:31,530
Then you need to investigate the root cause of that
133

133

00:04:31,530  -->  00:04:32,740
and solve those issues
134

134

00:04:32,740  -->  00:04:34,760
by bringing those metrics back to a normal level
135

135

00:04:34,760  -->  00:04:35,833
within your baseline.
136

136

00:04:37,130  -->  00:04:39,343
(upbeat music)
