1
00:00:00,090 --> 00:00:00,930
In this lesson,

2
00:00:00,930 --> 00:00:03,660
we're going to cover resilience and recovery testing.

3
00:00:03,660 --> 00:00:06,660
Now, resilience testing is used to assess a system's ability

4
00:00:06,660 --> 00:00:09,000
to withstand and adapt to disruptive events.

5
00:00:09,000 --> 00:00:10,620
Recovery testing, on the other hand,

6
00:00:10,620 --> 00:00:12,180
will evaluate a system's capacity

7
00:00:12,180 --> 00:00:15,300
to restore normal operations after a disruptive event.

8
00:00:15,300 --> 00:00:17,310
Now, at their core, these types of resilience

9
00:00:17,310 --> 00:00:19,950
and recovery testing will serve as a form of fire drill

10
00:00:19,950 --> 00:00:21,300
for our enterprise networks

11
00:00:21,300 --> 00:00:23,370
and our organization's operations.

12
00:00:23,370 --> 00:00:26,220
These tests can occur using either a tabletop exercise,

13
00:00:26,220 --> 00:00:28,290
a failover test, a simulation,

14
00:00:28,290 --> 00:00:31,170
or parallel processing to determine how your organization

15
00:00:31,170 --> 00:00:33,300
can survive a potentially disruptive event,

16
00:00:33,300 --> 00:00:35,430
like a power loss, a natural disaster,

17
00:00:35,430 --> 00:00:38,130
a ransomware attack, or even a data breach.

18
00:00:38,130 --> 00:00:41,310
Now first, we need to talk about tabletop exercises.

19
00:00:41,310 --> 00:00:44,580
A tabletop exercise is a simulated scenario-based discussion

20
00:00:44,580 --> 00:00:46,620
among your key stakeholders to assess

21
00:00:46,620 --> 00:00:48,450
and improve in organization's preparedness

22
00:00:48,450 --> 00:00:50,520
and response to a specific emergency

23
00:00:50,520 --> 00:00:52,650
or crisis situation without the need

24
00:00:52,650 --> 00:00:55,020
for the actual deployment of its resources.

25
00:00:55,020 --> 00:00:57,150
So what does a tabletop event look like?

26
00:00:57,150 --> 00:00:59,820
Well, it's kind of like storytelling with a twist.

27
00:00:59,820 --> 00:01:01,200
Imagine that you're sitting around a table

28
00:01:01,200 --> 00:01:03,900
with a bunch of other key stakeholders in your organization,

29
00:01:03,900 --> 00:01:06,300
and each of you is handed a basic script.

30
00:01:06,300 --> 00:01:08,190
Then a facilitator walks into the room

31
00:01:08,190 --> 00:01:10,110
and declares that something bad just happened.

32
00:01:10,110 --> 00:01:11,250
For example, they might say

33
00:01:11,250 --> 00:01:13,380
that your security operation center just detected

34
00:01:13,380 --> 00:01:15,600
that one of your domain controllers has been compromised

35
00:01:15,600 --> 00:01:16,860
by a nation-state actor,

36
00:01:16,860 --> 00:01:17,940
and they're then going to ask you

37
00:01:17,940 --> 00:01:19,710
what your organization is going to do

38
00:01:19,710 --> 00:01:21,300
based upon that piece of information

39
00:01:21,300 --> 00:01:22,950
about the suspected attack,

40
00:01:22,950 --> 00:01:25,230
which we call an exercise inject.

41
00:01:25,230 --> 00:01:26,160
At this point,

42
00:01:26,160 --> 00:01:28,200
all the stakeholders in the room are now thrown

43
00:01:28,200 --> 00:01:30,000
into this hypothetical disaster,

44
00:01:30,000 --> 00:01:31,260
and they're going to take turns discussing

45
00:01:31,260 --> 00:01:33,810
what actions each of their teams is going to conduct

46
00:01:33,810 --> 00:01:36,510
in response to this detected malicious activity.

47
00:01:36,510 --> 00:01:38,910
Now, the great thing about a tabletop exercise is

48
00:01:38,910 --> 00:01:40,080
that it lets each stakeholder

49
00:01:40,080 --> 00:01:41,250
and their team figure out

50
00:01:41,250 --> 00:01:43,890
how they're going to respond effectively to the given inject,

51
00:01:43,890 --> 00:01:45,990
and this is a fairly low-cost option to use

52
00:01:45,990 --> 00:01:48,150
while still providing an extremely engaging environment

53
00:01:48,150 --> 00:01:50,820
for those involved in this tabletop exercise.

54
00:01:50,820 --> 00:01:53,310
Now, as each stakeholder lays out their planned actions,

55
00:01:53,310 --> 00:01:55,410
the other stakeholders can identify gaps

56
00:01:55,410 --> 00:01:57,810
or seams in that plan that could cause issues

57
00:01:57,810 --> 00:01:59,430
during a real response too.

58
00:01:59,430 --> 00:02:02,460
So these plans will then get adjusted and captured as part

59
00:02:02,460 --> 00:02:05,040
of your organization's future standard operating procedures

60
00:02:05,040 --> 00:02:07,230
and your playbooks for a future incident response

61
00:02:07,230 --> 00:02:08,460
that you can use.

62
00:02:08,460 --> 00:02:11,250
Tabletops are also considered to be a form of team building

63
00:02:11,250 --> 00:02:12,420
because all the stakeholders

64
00:02:12,420 --> 00:02:14,310
and their teams are going to work together

65
00:02:14,310 --> 00:02:16,350
to try and solve the hypothetical disaster

66
00:02:16,350 --> 00:02:17,550
or malicious activities

67
00:02:17,550 --> 00:02:20,340
that are being injected into that given scenario.

68
00:02:20,340 --> 00:02:22,380
Second, we have a failover test.

69
00:02:22,380 --> 00:02:24,750
Now, a failover test is a controlled experiment

70
00:02:24,750 --> 00:02:26,760
that's designed to verify the seamless transition

71
00:02:26,760 --> 00:02:28,230
of a system or application

72
00:02:28,230 --> 00:02:31,380
from a primary component to a backup or secondary component

73
00:02:31,380 --> 00:02:32,640
in the event of a failure

74
00:02:32,640 --> 00:02:34,950
to ensure that uninterrupted functionality can be achieved

75
00:02:34,950 --> 00:02:37,050
during a disaster or other incident.

76
00:02:37,050 --> 00:02:39,870
For example, if your organization has already decided

77
00:02:39,870 --> 00:02:41,520
that if you have a large-scale disaster

78
00:02:41,520 --> 00:02:42,960
at your East Coast Data Center,

79
00:02:42,960 --> 00:02:44,370
that you're going to shift all your operations

80
00:02:44,370 --> 00:02:45,720
to an alternative hot site

81
00:02:45,720 --> 00:02:47,220
that you've located on the West Coast

82
00:02:47,220 --> 00:02:48,390
in your data center,

83
00:02:48,390 --> 00:02:50,670
this can actually be done using a failover test,

84
00:02:50,670 --> 00:02:52,740
where we can actually attempt to do this cutover

85
00:02:52,740 --> 00:02:54,900
from the East Coast to the West Coast.

86
00:02:54,900 --> 00:02:57,390
Now, since we're actually going to be taking all these actions,

87
00:02:57,390 --> 00:02:59,040
this does involve more resources,

88
00:02:59,040 --> 00:03:01,350
more time, and more energy to perform.

89
00:03:01,350 --> 00:03:03,240
But by performing a failover test,

90
00:03:03,240 --> 00:03:04,230
you're going to be able to verify

91
00:03:04,230 --> 00:03:06,270
that your planned actions will actually work

92
00:03:06,270 --> 00:03:08,100
when a real disaster strikes.

93
00:03:08,100 --> 00:03:09,540
For example, I used to work

94
00:03:09,540 --> 00:03:12,540
for a large organization located near Washington, D.C.,

95
00:03:12,540 --> 00:03:14,940
and if our facility was affected by a disaster,

96
00:03:14,940 --> 00:03:17,010
we had a plan where we could shift our operations

97
00:03:17,010 --> 00:03:20,670
to an alternative hot site located about 1,500 miles away.

98
00:03:20,670 --> 00:03:22,980
To ensure we could execute this plan during an emergency,

99
00:03:22,980 --> 00:03:25,170
we actually tested this using a failover test

100
00:03:25,170 --> 00:03:27,060
at least one time per year.

101
00:03:27,060 --> 00:03:28,740
In preparation for the failover test,

102
00:03:28,740 --> 00:03:31,050
we actually would fly out a small team to the hot site

103
00:03:31,050 --> 00:03:33,270
to ensure our operations could continue smoothly.

104
00:03:33,270 --> 00:03:36,360
After any issues, we always had in place a rollback plan,

105
00:03:36,360 --> 00:03:38,850
where we could shift operations back to our main facility

106
00:03:38,850 --> 00:03:39,683
while we figured out

107
00:03:39,683 --> 00:03:41,820
why our failover didn't work out as expected,

108
00:03:41,820 --> 00:03:43,410
which only happened once or twice

109
00:03:43,410 --> 00:03:45,420
over the several years I was there.

110
00:03:45,420 --> 00:03:47,190
Third, we have simulations.

111
00:03:47,190 --> 00:03:49,110
Now, a simulation is a computer generated

112
00:03:49,110 --> 00:03:51,810
or artificial representation of a real-world system

113
00:03:51,810 --> 00:03:54,360
or scenario that's used to mimic various conditions,

114
00:03:54,360 --> 00:03:56,880
interactions, or failures, to assess how a system

115
00:03:56,880 --> 00:04:00,210
or process will respond in order to refine its performance.

116
00:04:00,210 --> 00:04:01,860
Now, unlike a tabletop exercise

117
00:04:01,860 --> 00:04:03,570
where the threat is simply a scenario presented

118
00:04:03,570 --> 00:04:04,680
on a piece of paper,

119
00:04:04,680 --> 00:04:07,230
a simulation actually allows your incident responders

120
00:04:07,230 --> 00:04:08,370
and system administrators

121
00:04:08,370 --> 00:04:11,100
to perform their response actions inside of a virtualized

122
00:04:11,100 --> 00:04:12,750
or simulated environment.

123
00:04:12,750 --> 00:04:15,510
For example, with our modern cloud-based infrastructures,

124
00:04:15,510 --> 00:04:16,860
we can actually spin up a virtual

125
00:04:16,860 --> 00:04:18,750
or simulated version of our corporate network

126
00:04:18,750 --> 00:04:19,950
inside of the cloud,

127
00:04:19,950 --> 00:04:22,740
and then we can have a red team attack that network,

128
00:04:22,740 --> 00:04:24,750
while our defenders, who are known as the blue team,

129
00:04:24,750 --> 00:04:26,880
are trying to detect that red team's attacks

130
00:04:26,880 --> 00:04:29,160
and utilize their proper incident response techniques

131
00:04:29,160 --> 00:04:31,200
to isolate the attackers from the network

132
00:04:31,200 --> 00:04:34,620
and remove any infected systems from that simulated network.

133
00:04:34,620 --> 00:04:37,200
With a simulation, your personnel are able to respond

134
00:04:37,200 --> 00:04:39,240
to the simulated crisis in real time,

135
00:04:39,240 --> 00:04:40,440
which helps test your reactions

136
00:04:40,440 --> 00:04:41,970
in a more realistic environment

137
00:04:41,970 --> 00:04:44,670
that you can then evaluate not just the plan itself

138
00:04:44,670 --> 00:04:46,020
but also your staff members

139
00:04:46,020 --> 00:04:47,250
and their ability to work

140
00:04:47,250 --> 00:04:50,250
during the unpredictable chaos of a cyber attack.

141
00:04:50,250 --> 00:04:51,750
After the simulation is over,

142
00:04:51,750 --> 00:04:53,730
all the observers are going to provide feedback

143
00:04:53,730 --> 00:04:56,430
on the attacker's performance and the defender's performance

144
00:04:56,430 --> 00:04:58,440
so that each side can learn from their mistakes

145
00:04:58,440 --> 00:05:00,090
and improve their skills.

146
00:05:00,090 --> 00:05:02,940
Now, fourth and finally, we have parallel processing.

147
00:05:02,940 --> 00:05:05,310
Parallel processing involves replicating your data

148
00:05:05,310 --> 00:05:08,160
and systems processes onto a secondary system,

149
00:05:08,160 --> 00:05:09,960
and then you run both the primary

150
00:05:09,960 --> 00:05:13,290
and secondary systems together at the same time.

151
00:05:13,290 --> 00:05:15,360
Our goal when we conduct parallel processing

152
00:05:15,360 --> 00:05:16,590
is to check the reliability

153
00:05:16,590 --> 00:05:18,690
and stability of your secondary setup

154
00:05:18,690 --> 00:05:20,400
to make sure it can handle processing data

155
00:05:20,400 --> 00:05:22,470
without disrupting your day-to-day operations

156
00:05:22,470 --> 00:05:24,060
during a disaster.

157
00:05:24,060 --> 00:05:26,970
Now, parallel processing does require meticulous planning,

158
00:05:26,970 --> 00:05:29,640
flawless execution, and an eagle eye for detail

159
00:05:29,640 --> 00:05:31,110
to ensure that there is zero disruption

160
00:05:31,110 --> 00:05:32,580
to your ongoing operations.

161
00:05:32,580 --> 00:05:35,370
So it is something you have to do very carefully.

162
00:05:35,370 --> 00:05:38,520
Now, in resilience testing, we often use parallel processing

163
00:05:38,520 --> 00:05:40,110
to help test the system's ability

164
00:05:40,110 --> 00:05:42,570
to handle multiple failure scenarios simultaneously,

165
00:05:42,570 --> 00:05:43,770
such as network outages,

166
00:05:43,770 --> 00:05:45,900
hardware failures, or data corruption.

167
00:05:45,900 --> 00:05:47,430
When it comes to recovery testing,

168
00:05:47,430 --> 00:05:48,870
we use parallel processing

169
00:05:48,870 --> 00:05:51,030
to evaluate how efficiently a system can recover

170
00:05:51,030 --> 00:05:52,770
from multiple failure points.

171
00:05:52,770 --> 00:05:55,380
So remember, when it comes to tabletop exercises,

172
00:05:55,380 --> 00:05:58,440
failover tests, simulations, and parallel processing,

173
00:05:58,440 --> 00:06:00,420
our goal here is to plan for the worst

174
00:06:00,420 --> 00:06:02,130
and learn how to overcome any obstacles

175
00:06:02,130 --> 00:06:03,930
that we might potentially face.

176
00:06:03,930 --> 00:06:05,430
Resilience and recovery testing

177
00:06:05,430 --> 00:06:07,200
is not just about creating a plan,

178
00:06:07,200 --> 00:06:09,690
but it's about creating a plan, testing that plan,

179
00:06:09,690 --> 00:06:11,730
and assuring ourselves and our stakeholders

180
00:06:11,730 --> 00:06:13,890
that our plan will really work when a disaster

181
00:06:13,890 --> 00:06:16,200
or incident does eventually occur to us.

182
00:06:16,200 --> 00:06:18,840
Each of these methods, including tabletop exercises,

183
00:06:18,840 --> 00:06:21,570
failovers, simulations, and parallel processing,

184
00:06:21,570 --> 00:06:23,490
are going to offer us with different insights,

185
00:06:23,490 --> 00:06:25,440
the ability to pinpoint various weaknesses,

186
00:06:25,440 --> 00:06:27,360
and they'll help us to prepare our teams

187
00:06:27,360 --> 00:06:29,160
for just about anything that an attacker

188
00:06:29,160 --> 00:06:31,110
or the environment might throw at us.

189
00:06:31,110 --> 00:06:32,490
With all these tests, though,

190
00:06:32,490 --> 00:06:34,590
this isn't going to be a one-time thing.

191
00:06:34,590 --> 00:06:36,870
Instead, we have to keep doing it routinely.

192
00:06:36,870 --> 00:06:39,720
You should regularly test, adapt, and improve your plans

193
00:06:39,720 --> 00:06:41,130
to ensure that your plans are evolving

194
00:06:41,130 --> 00:06:43,590
in response to the rapidly changing threat environment

195
00:06:43,590 --> 00:06:46,530
and any new vulnerabilities that may have emerged over time.

196
00:06:46,530 --> 00:06:48,630
After all, your organization's resilience

197
00:06:48,630 --> 00:06:50,160
is not just a check in the box.

198
00:06:50,160 --> 00:06:52,590
It is an ongoing, ever-continuing journey

199
00:06:52,590 --> 00:06:54,090
towards continual improvement.

