1
00:00:00,600 --> 00:00:02,130
In this lesson, we're going to discuss

2
00:00:02,130 --> 00:00:03,600
Root Cause Analysis.

3
00:00:03,600 --> 00:00:06,630
Now, a root cause analysis is a systematic process

4
00:00:06,630 --> 00:00:09,327
to identify the initial source of the incident

5
00:00:09,327 --> 00:00:11,670
and to prevent it from occurring again.

6
00:00:11,670 --> 00:00:14,010
This analysis is usually going to occur

7
00:00:14,010 --> 00:00:15,930
using a four-step process.

8
00:00:15,930 --> 00:00:18,510
First, we define the scope of the incident.

9
00:00:18,510 --> 00:00:21,000
Second, we determine the causal relationships

10
00:00:21,000 --> 00:00:22,560
that led to that incident.

11
00:00:22,560 --> 00:00:24,990
Third, we identify an effective solution.

12
00:00:24,990 --> 00:00:26,370
And fourth, we implement

13
00:00:26,370 --> 00:00:27,510
and track the solution

14
00:00:27,510 --> 00:00:30,600
to ensure that the incident is fully resolved.

15
00:00:30,600 --> 00:00:32,790
So let's go through a quick example.

16
00:00:32,790 --> 00:00:34,380
Let's say you had a malware infection

17
00:00:34,380 --> 00:00:35,700
inside your organization.

18
00:00:35,700 --> 00:00:38,010
First, you need to determine the initial cause

19
00:00:38,010 --> 00:00:39,690
and scope of that incident.

20
00:00:39,690 --> 00:00:40,830
Maybe it was caused

21
00:00:40,830 --> 00:00:43,080
because malware was introduced into your network

22
00:00:43,080 --> 00:00:45,540
because somebody plugged in a USB thumb drive

23
00:00:45,540 --> 00:00:46,590
into their workstation,

24
00:00:46,590 --> 00:00:49,440
or they clicked the link in a spear-phishing campaign,

25
00:00:49,440 --> 00:00:51,390
or they visited a malicious website.

26
00:00:51,390 --> 00:00:54,180
Now, whatever the initial attack vector was,

27
00:00:54,180 --> 00:00:55,860
our goal is to identify that

28
00:00:55,860 --> 00:00:58,500
so we can prevent it from happening again.

29
00:00:58,500 --> 00:01:00,480
If somebody caused this incident to occur

30
00:01:00,480 --> 00:01:03,840
by plugging in a thumb drive that was infected with malware,

31
00:01:03,840 --> 00:01:06,060
there are a couple ways that we can prevent this.

32
00:01:06,060 --> 00:01:08,490
One way would be to ensure that all of our workstations

33
00:01:08,490 --> 00:01:11,730
have the latest version of antivirus installed on them.

34
00:01:11,730 --> 00:01:14,820
This way, anytime you insert a disc, a CD,

35
00:01:14,820 --> 00:01:19,050
or a DVD, download a file or plug in a USB thumb drive,

36
00:01:19,050 --> 00:01:20,880
it's going to be scanned for viruses

37
00:01:20,880 --> 00:01:22,920
and other malware before allowing you

38
00:01:22,920 --> 00:01:25,350
to read from the device.

39
00:01:25,350 --> 00:01:26,550
Another thing you might do

40
00:01:26,550 --> 00:01:27,930
is to prevent the data transfer

41
00:01:27,930 --> 00:01:30,180
from USB devices like thumb drives

42
00:01:30,180 --> 00:01:31,440
and external hard drives

43
00:01:31,440 --> 00:01:34,260
for all users on your enterprise network in the future,

44
00:01:34,260 --> 00:01:35,550
this will be another great way

45
00:01:35,550 --> 00:01:38,160
to stop data from getting from this infected device

46
00:01:38,160 --> 00:01:39,780
onto your system.

47
00:01:39,780 --> 00:01:42,510
Now, in addition to that, we might also find out

48
00:01:42,510 --> 00:01:44,340
that this particular piece of malware

49
00:01:44,340 --> 00:01:46,740
was only infected against certain types of machines

50
00:01:46,740 --> 00:01:48,510
that were running a certain version of windows

51
00:01:48,510 --> 00:01:51,390
or we're missing some type of security patch.

52
00:01:51,390 --> 00:01:53,760
And in those cases, we would identify that weakness,

53
00:01:53,760 --> 00:01:55,380
and then we will create a solution

54
00:01:55,380 --> 00:01:58,350
to identify those things such as upgrading the version

55
00:01:58,350 --> 00:02:00,720
of Windows or installing the security patches,

56
00:02:00,720 --> 00:02:03,750
as that will protect us from this vulnerability.

57
00:02:03,750 --> 00:02:05,790
So now that we've gathered this information,

58
00:02:05,790 --> 00:02:07,860
we have to define and scope our incident.

59
00:02:07,860 --> 00:02:10,289
We should know how many machines have been affected

60
00:02:10,289 --> 00:02:13,080
and how many users have been affected by this incident,

61
00:02:13,080 --> 00:02:16,320
and what operational impacts that has.

62
00:02:16,320 --> 00:02:19,290
Then we go and we determine the causal relationships

63
00:02:19,290 --> 00:02:20,850
that led to that incident.

64
00:02:20,850 --> 00:02:22,770
In this case, somebody installed malware

65
00:02:22,770 --> 00:02:24,420
using a USB thumb drive.

66
00:02:24,420 --> 00:02:28,257
So we need to identify an effective solution to stop that,

67
00:02:28,257 --> 00:02:30,630
and we came up with a handful here,

68
00:02:30,630 --> 00:02:32,520
just as we were going through this example.

69
00:02:32,520 --> 00:02:34,440
Things like adding antivirus

70
00:02:34,440 --> 00:02:36,300
or preventing data from being read

71
00:02:36,300 --> 00:02:38,370
from a USB mass storage device,

72
00:02:38,370 --> 00:02:40,950
or installing a newer version of Windows,

73
00:02:40,950 --> 00:02:43,050
or even an updated security patch

74
00:02:43,050 --> 00:02:45,270
for that particular vulnerability.

75
00:02:45,270 --> 00:02:47,580
And then this brings us to the fourth step,

76
00:02:47,580 --> 00:02:48,480
which is to implement

77
00:02:48,480 --> 00:02:49,950
and track the solutions to ensure

78
00:02:49,950 --> 00:02:51,720
that the incident is fully handled.

79
00:02:51,720 --> 00:02:54,720
Now, in this case, we might have a single system

80
00:02:54,720 --> 00:02:56,490
that was infected by a piece of malware

81
00:02:56,490 --> 00:02:58,470
on the USB device.

82
00:02:58,470 --> 00:03:00,660
So we've identified what it was,

83
00:03:00,660 --> 00:03:03,300
we determined the relationship that led to the incident,

84
00:03:03,300 --> 00:03:06,450
and we've identified a couple of effective solutions.

85
00:03:06,450 --> 00:03:08,850
In this case, I'm going to want to go ahead

86
00:03:08,850 --> 00:03:11,580
and ensure that I block mass storage devices

87
00:03:11,580 --> 00:03:13,440
from being read on the user system,

88
00:03:13,440 --> 00:03:15,750
and I want to ensure that their antivirus

89
00:03:15,750 --> 00:03:17,883
and antimalware solution is up to date.

90
00:03:18,750 --> 00:03:22,050
And now I need to implement and track those action items.

91
00:03:22,050 --> 00:03:24,720
In this case, I would tell my system administrators

92
00:03:24,720 --> 00:03:26,520
through a change management process

93
00:03:26,520 --> 00:03:28,260
to install the registry change

94
00:03:28,260 --> 00:03:30,150
that will prevent USB devices

95
00:03:30,150 --> 00:03:31,920
from being read on that machine

96
00:03:31,920 --> 00:03:34,410
and make sure that it is part of our asset management

97
00:03:34,410 --> 00:03:37,050
and change management process to update the fact

98
00:03:37,050 --> 00:03:38,790
that machine has been updated

99
00:03:38,790 --> 00:03:40,500
to the latest version of Windows

100
00:03:40,500 --> 00:03:42,660
and has the latest security patches.

101
00:03:42,660 --> 00:03:44,130
By doing all of these things,

102
00:03:44,130 --> 00:03:45,840
we have now contained this incident

103
00:03:45,840 --> 00:03:47,310
and figured out the root cause,

104
00:03:47,310 --> 00:03:49,440
and stopped it from happening again.

105
00:03:49,440 --> 00:03:52,020
Now, I like to take this one step further.

106
00:03:52,020 --> 00:03:54,540
Because we have identified it on this single machine,

107
00:03:54,540 --> 00:03:56,730
we also may need to look across the network

108
00:03:56,730 --> 00:03:58,380
and see if there are other machines

109
00:03:58,380 --> 00:04:00,090
that could have been affected as well.

110
00:04:00,090 --> 00:04:02,190
For example, if we figure out the reason

111
00:04:02,190 --> 00:04:04,530
that this piece of malware was able to be installed

112
00:04:04,530 --> 00:04:07,470
is because you're running Windows 10 instead of Windows 11,

113
00:04:07,470 --> 00:04:09,540
then we might need to look at the rest of our network

114
00:04:09,540 --> 00:04:13,200
and say, "Hmm, how many Windows 10 machines do I have?"

115
00:04:13,200 --> 00:04:16,050
If there are a lot of them, it may be a very large solution

116
00:04:16,050 --> 00:04:18,810
for us to try to upgrade all systems to Windows 11

117
00:04:18,810 --> 00:04:21,329
and prevent this vulnerability from being exploited again,

118
00:04:21,329 --> 00:04:24,210
that's the real benefit of doing a root cause analysis.

119
00:04:24,210 --> 00:04:26,760
You need to figure out what caused the incident

120
00:04:26,760 --> 00:04:29,280
and then see how many other things across your network

121
00:04:29,280 --> 00:04:31,770
or across your organization are going to have to have

122
00:04:31,770 --> 00:04:34,290
that same feature set that could then be vulnerable

123
00:04:34,290 --> 00:04:35,700
to the same type of attacks

124
00:04:35,700 --> 00:04:38,490
so you can mitigate those vulnerabilities.

125
00:04:38,490 --> 00:04:40,860
Now, when you conduct a root cause analysis,

126
00:04:40,860 --> 00:04:42,990
you must make sure that this process uses

127
00:04:42,990 --> 00:04:44,520
a no-blame approach.

128
00:04:44,520 --> 00:04:46,830
In other words, the root cause analysis process

129
00:04:46,830 --> 00:04:48,930
is not designed to point fingers or assign blame

130
00:04:48,930 --> 00:04:50,460
to an individual or team.

131
00:04:50,460 --> 00:04:52,500
Instead, our primary goal is to identify

132
00:04:52,500 --> 00:04:54,030
the root cause of the incident

133
00:04:54,030 --> 00:04:56,490
and develop effective preventative measures

134
00:04:56,490 --> 00:04:59,940
to prevent the incident from occurring again in the future.

135
00:04:59,940 --> 00:05:02,550
Using a no-blame approach can encourage open

136
00:05:02,550 --> 00:05:03,720
and honest reporting,

137
00:05:03,720 --> 00:05:06,000
which is crucial for improving cybersecurity practices

138
00:05:06,000 --> 00:05:07,440
within an organization.

139
00:05:07,440 --> 00:05:09,210
Consider the real-world example

140
00:05:09,210 --> 00:05:13,440
of the two plane crash involving a Boeing 737 MAX

141
00:05:13,440 --> 00:05:17,940
that occurred back in October 2018 and March of 2019.

142
00:05:17,940 --> 00:05:18,990
After the crashes,

143
00:05:18,990 --> 00:05:21,510
the United States National Traffic Safety Board,

144
00:05:21,510 --> 00:05:24,570
known as the NTSB, joined the investigation

145
00:05:24,570 --> 00:05:26,670
to conduct a root cause analysis.

146
00:05:26,670 --> 00:05:27,930
During their investigation,

147
00:05:27,930 --> 00:05:30,420
the NTSB conducted an independent,

148
00:05:30,420 --> 00:05:32,610
no-blame root cause analysis

149
00:05:32,610 --> 00:05:34,680
with the primary objective of determining

150
00:05:34,680 --> 00:05:36,540
what went wrong and why.

151
00:05:36,540 --> 00:05:39,030
Their investigations are meticulous,

152
00:05:39,030 --> 00:05:40,620
thorough, and independent,

153
00:05:40,620 --> 00:05:42,480
and they often evolve multiple experts

154
00:05:42,480 --> 00:05:44,070
across various fields.

155
00:05:44,070 --> 00:05:47,220
For instance, NTSB might look at an airline crash

156
00:05:47,220 --> 00:05:49,560
to see if it was caused by one or more factors,

157
00:05:49,560 --> 00:05:52,110
including pilot error, weather conditions,

158
00:05:52,110 --> 00:05:53,670
or technical malfunction.

159
00:05:53,670 --> 00:05:55,830
NTSB's investigators do not focus

160
00:05:55,830 --> 00:05:57,900
solely on blaming the pilot, but instead,

161
00:05:57,900 --> 00:05:59,940
they analyzed the entire chain of events,

162
00:05:59,940 --> 00:06:03,000
including weather data, aircraft maintenance records,

163
00:06:03,000 --> 00:06:06,330
communication transcripts, and cockpit voice recordings.

164
00:06:06,330 --> 00:06:10,327
Now, in the case of that Boeing 737-MAX crashes of 2018

165
00:06:11,257 --> 00:06:14,370
and 2019, it was determined that both crashes were caused

166
00:06:14,370 --> 00:06:17,460
by issues in aircraft's new flight control software

167
00:06:17,460 --> 00:06:20,100
that repeatedly pushed the jet's nose down

168
00:06:20,100 --> 00:06:22,770
due to a defective sensor on the aircraft.

169
00:06:22,770 --> 00:06:25,380
This essentially removed control from the pilots

170
00:06:25,380 --> 00:06:27,750
and caused both planes to crash into the ground

171
00:06:27,750 --> 00:06:29,220
since the pilot could not overcome

172
00:06:29,220 --> 00:06:32,490
the new yet defective maneuvering characteristic

173
00:06:32,490 --> 00:06:35,550
augmentation system used as a flight control system

174
00:06:35,550 --> 00:06:37,200
on each planes.

175
00:06:37,200 --> 00:06:40,350
So why does the NTSB use a no-blame approach

176
00:06:40,350 --> 00:06:42,510
when conducting their root cause analysis?

177
00:06:42,510 --> 00:06:45,690
Well, the NTSB is known for recognizing

178
00:06:45,690 --> 00:06:48,030
that human errors are often the result

179
00:06:48,030 --> 00:06:50,940
of systemic issues within the aviation industry,

180
00:06:50,940 --> 00:06:53,610
such as training procedures, equipment design,

181
00:06:53,610 --> 00:06:55,830
or regulatory oversight.

182
00:06:55,830 --> 00:06:57,660
By using a no-blame approach,

183
00:06:57,660 --> 00:07:00,810
the NTSB can identify these systemic weaknesses

184
00:07:00,810 --> 00:07:02,280
and recommend improvements

185
00:07:02,280 --> 00:07:05,070
to prevent similar incidents in the future.

186
00:07:05,070 --> 00:07:07,440
In the case of the Boeing 737-MAX,

187
00:07:07,440 --> 00:07:10,710
there are almost 400 planes in operation around the world.

188
00:07:10,710 --> 00:07:13,620
So the NTSB and other country regulators

189
00:07:13,620 --> 00:07:15,060
quickly grounded these planes

190
00:07:15,060 --> 00:07:17,160
until the root cause was identified,

191
00:07:17,160 --> 00:07:20,010
and the defective flight control software was reviewed

192
00:07:20,010 --> 00:07:22,830
line by line to ensure that no other vulnerabilities

193
00:07:22,830 --> 00:07:26,640
could be identified that might cause future crashes to occur.

194
00:07:26,640 --> 00:07:29,310
So remember, when you conduct your root cause analysis,

195
00:07:29,310 --> 00:07:32,040
remember that our primary purpose is not to assign blame,

196
00:07:32,040 --> 00:07:34,410
but instead to figure out exactly what has occurred

197
00:07:34,410 --> 00:07:36,570
to cause the incident in the first place.

198
00:07:36,570 --> 00:07:39,390
By conducting a root cause analysis without blame,

199
00:07:39,390 --> 00:07:40,920
we can better identify vulnerabilities

200
00:07:40,920 --> 00:07:43,110
or weaknesses in our security practices

201
00:07:43,110 --> 00:07:46,290
and create more robust protections against cyber threats.

202
00:07:46,290 --> 00:07:48,210
Creating this no-blame culture

203
00:07:48,210 --> 00:07:51,150
can also encourage collaboration, transparency,

204
00:07:51,150 --> 00:07:54,420
and focus on solutions rather than assigning fault

205
00:07:54,420 --> 00:07:56,613
during or after an incident response.

