1
00:00:06,586 --> 00:00:08,610
- In this lesson 13.2,

2
00:00:08,610 --> 00:00:12,330
we're gonna focus in
on resiliency concepts.

3
00:00:12,330 --> 00:00:15,540
Now, resiliency is the
capability to continue operating

4
00:00:15,540 --> 00:00:17,520
even when there's been a disruption

5
00:00:17,520 --> 00:00:20,223
or there's abnormal operating conditions.

6
00:00:21,960 --> 00:00:25,710
I wanna make sure we distinguish
resiliency from redundancy.

7
00:00:25,710 --> 00:00:28,950
Redundancy is when we
duplicate critical components

8
00:00:28,950 --> 00:00:32,910
or functions with the intention
of increasing reliability

9
00:00:32,910 --> 00:00:34,890
and mitigating the risks
that are associated

10
00:00:34,890 --> 00:00:37,563
with a single point of
failure known as an SPOF.

11
00:00:38,430 --> 00:00:40,770
Now, fault tolerance is a capability

12
00:00:40,770 --> 00:00:42,840
of a system to continue to operate

13
00:00:42,840 --> 00:00:44,490
even in the event of failure

14
00:00:44,490 --> 00:00:47,310
of one or more system components.

15
00:00:47,310 --> 00:00:49,590
Now, categories of resiliency,

16
00:00:49,590 --> 00:00:51,900
are going to include system storage,

17
00:00:51,900 --> 00:00:55,260
power transmission, and site.

18
00:00:55,260 --> 00:00:56,910
So let's go look at all of these.

19
00:00:58,620 --> 00:01:00,870
So for system resiliency itself, right?

20
00:01:00,870 --> 00:01:03,030
Keeping our systems operating

21
00:01:03,030 --> 00:01:05,640
some options we have are load balancing,

22
00:01:05,640 --> 00:01:10,620
clustering, high
availability, and fail secure.

23
00:01:10,620 --> 00:01:12,600
So load balancing involves distributing

24
00:01:12,600 --> 00:01:14,010
incoming network traffic

25
00:01:14,010 --> 00:01:16,440
across multiple independent systems

26
00:01:16,440 --> 00:01:18,720
to ensure that no single server

27
00:01:18,720 --> 00:01:20,763
becomes overwhelmed with requests.

28
00:01:22,470 --> 00:01:25,320
Clustering involves grouping
multiple systems together

29
00:01:25,320 --> 00:01:28,170
to form a single logical unit or cluster

30
00:01:28,170 --> 00:01:30,720
and we can lose individual
components to that cluster

31
00:01:30,720 --> 00:01:32,640
and still be operating.

32
00:01:32,640 --> 00:01:35,010
High availability, you'll
see abbreviate is HA,

33
00:01:35,010 --> 00:01:37,590
is the automatic failover capability,

34
00:01:37,590 --> 00:01:40,500
which reduces or eliminates
the need to activate

35
00:01:40,500 --> 00:01:41,460
redundant hardware.

36
00:01:41,460 --> 00:01:45,000
Now, there are two approaches,
asymmetric and symmetric.

37
00:01:45,000 --> 00:01:46,050
In asymmetric,

38
00:01:46,050 --> 00:01:48,510
we have what's known as
an active passive pair

39
00:01:48,510 --> 00:01:50,850
where one device is active
and the other device

40
00:01:50,850 --> 00:01:52,920
is powered on but not
really doing anything

41
00:01:52,920 --> 00:01:54,900
and it's busy listening for a heartbeat.

42
00:01:54,900 --> 00:01:56,970
Boom, boom, boom, boom, boom, right?

43
00:01:56,970 --> 00:01:59,400
Is that primary device operating?

44
00:01:59,400 --> 00:02:01,020
And if it doesn't hear that heartbeat

45
00:02:01,020 --> 00:02:04,530
at some point it says,
okay, I need to come online.

46
00:02:04,530 --> 00:02:07,170
In symmetric, we have an active pair

47
00:02:07,170 --> 00:02:09,960
where both devices are operational, right?

48
00:02:09,960 --> 00:02:13,263
And so one device fails, the
other device just continues on.

49
00:02:16,230 --> 00:02:19,380
And then we wanna think about
the principle of fail secure

50
00:02:19,380 --> 00:02:20,790
in terms of resiliency.

51
00:02:20,790 --> 00:02:23,460
The idea that a failure will always result

52
00:02:23,460 --> 00:02:26,013
in a secure or trustworthy state.

53
00:02:27,390 --> 00:02:30,600
The rate is redundant,
array of inexpensive discs

54
00:02:30,600 --> 00:02:34,500
which is a data storage
virtualization technology.

55
00:02:34,500 --> 00:02:35,580
Now, the idea of rate

56
00:02:35,580 --> 00:02:38,190
is that we combine multiple
disc drive components

57
00:02:38,190 --> 00:02:39,930
into one or more logical units

58
00:02:39,930 --> 00:02:43,350
for the purpose of fault
tolerance and data redundancy

59
00:02:43,350 --> 00:02:46,170
and or performance improvement.

60
00:02:46,170 --> 00:02:50,880
RAID can be configured for
mirroring, striping, or both.

61
00:02:50,880 --> 00:02:53,790
The disc mirroring is the
process of writing data

62
00:02:53,790 --> 00:02:56,220
on two partitions on separate discs

63
00:02:56,220 --> 00:02:58,807
where just striping is the
process of dividing data

64
00:02:58,807 --> 00:03:02,070
into blocks and then
spreading those data blocks

65
00:03:02,070 --> 00:03:04,443
across multiple storage devices.

66
00:03:05,640 --> 00:03:08,010
So let's take a look at RAID 1 and RAID 5

67
00:03:08,010 --> 00:03:11,820
which are really the two most
common configurations of RAID.

68
00:03:11,820 --> 00:03:14,340
In RAID 1, that's the top graphic you see,

69
00:03:14,340 --> 00:03:17,100
we have two drives, right?

70
00:03:17,100 --> 00:03:20,430
And everything that happens on disc zero

71
00:03:20,430 --> 00:03:22,650
is also gonna happen on disc one.

72
00:03:22,650 --> 00:03:23,910
So you can see on each drive

73
00:03:23,910 --> 00:03:28,350
we have 1A, 1A, 2A2A, 3A3A, 4A4A.

74
00:03:28,350 --> 00:03:29,430
So everything that's happening,

75
00:03:29,430 --> 00:03:30,750
everything that's being written to one

76
00:03:30,750 --> 00:03:32,580
is being written to the other.

77
00:03:32,580 --> 00:03:36,030
That means that we could lose
either one of those drives

78
00:03:36,030 --> 00:03:37,830
and continue to be operating.

79
00:03:37,830 --> 00:03:39,000
Now they're both active, right?

80
00:03:39,000 --> 00:03:40,740
It's an active, active pair,

81
00:03:40,740 --> 00:03:44,040
but you know lose either one
and we can continue to operate.

82
00:03:44,040 --> 00:03:45,690
So that's dismirroring.

83
00:03:45,690 --> 00:03:47,130
There's still a single point of failure

84
00:03:47,130 --> 00:03:48,540
though with dismirroring,

85
00:03:48,540 --> 00:03:52,890
and that is what if we
only have one controller?

86
00:03:52,890 --> 00:03:54,240
So if we have one controller

87
00:03:54,240 --> 00:03:56,790
we still an SPOF, a
single point of failure.

88
00:03:56,790 --> 00:03:58,230
And if that controller fails

89
00:03:58,230 --> 00:04:00,630
well we can't access either drive.

90
00:04:00,630 --> 00:04:02,280
So in some configurations

91
00:04:02,280 --> 00:04:04,530
we'll give each drive its own controller.

92
00:04:04,530 --> 00:04:07,533
And when that's true, we'll
call it disc duplexing.

93
00:04:08,910 --> 00:04:11,910
We come down to the lower
picture rate, level five,

94
00:04:11,910 --> 00:04:16,230
this is gonna be an
example of disc striping

95
00:04:16,230 --> 00:04:18,930
and we're doing disc
striping with a parody bit.

96
00:04:18,930 --> 00:04:20,430
So disc striping in of itself

97
00:04:20,430 --> 00:04:22,410
stripes the data across multiple drives

98
00:04:22,410 --> 00:04:26,160
but doesn't really give us any resiliency.

99
00:04:26,160 --> 00:04:29,490
But if we do a parody
bit across each stripe

100
00:04:29,490 --> 00:04:33,270
then we have the ability to
recover a drive if it fails.

101
00:04:33,270 --> 00:04:37,080
So you can see in the first
instance we've got 1A, 2A,

102
00:04:37,080 --> 00:04:38,910
3A, and then 4P.

103
00:04:38,910 --> 00:04:40,077
That's a parody bit.

104
00:04:40,077 --> 00:04:43,320
And so what's written right
here in that parody bit

105
00:04:43,320 --> 00:04:47,370
is enough information
that if we lost the first,

106
00:04:47,370 --> 00:04:49,860
second, or third drive
we could reconstruct

107
00:04:49,860 --> 00:04:53,340
the data that was on that
pass from that parody bit.

108
00:04:53,340 --> 00:04:55,560
And that parody bit's gonna move around.

109
00:04:55,560 --> 00:04:59,371
So you can see on the next
pass we have 1B, 2B, 3P

110
00:04:59,371 --> 00:05:02,670
that's the parody and then 4B.

111
00:05:02,670 --> 00:05:07,670
And on the next one we have
1C, two parody, 3C, 4C.

112
00:05:07,920 --> 00:05:09,390
So our goal in RAID 5

113
00:05:09,390 --> 00:05:11,130
is that we have a performance improvement

114
00:05:11,130 --> 00:05:15,750
because we're striping, but we
have the availability, right?

115
00:05:15,750 --> 00:05:18,450
The redundancy, the
resiliency, if you will

116
00:05:18,450 --> 00:05:20,670
because we can lose
any one of those drives

117
00:05:20,670 --> 00:05:22,050
and we can continue operating

118
00:05:22,050 --> 00:05:23,610
because the data that was on that drive

119
00:05:23,610 --> 00:05:26,343
can be reconstructed from the parody bit.

120
00:05:28,980 --> 00:05:30,960
Next up is power resiliency.

121
00:05:30,960 --> 00:05:34,560
So we can have redundancy
having a UPS battery backup,

122
00:05:34,560 --> 00:05:37,800
having a generator,
and supplier diversity.

123
00:05:37,800 --> 00:05:40,920
So redundancy just means
at the component level

124
00:05:40,920 --> 00:05:44,100
having two or more power supplies or fans

125
00:05:44,100 --> 00:05:47,340
as power supplies are fans,
they break, they go, right?

126
00:05:47,340 --> 00:05:50,070
So by making sure that we have
multiple of those in a system

127
00:05:50,070 --> 00:05:51,063
we keep operating.

128
00:05:52,050 --> 00:05:55,170
But what if we actually
lose our power source?

129
00:05:55,170 --> 00:05:58,140
Well, we can use a UPS
uninterruptible power supplies

130
00:05:58,140 --> 00:06:00,300
and that's referred to
as a battery backup.

131
00:06:00,300 --> 00:06:03,390
And uninterruptible power
supply provides backup power

132
00:06:03,390 --> 00:06:05,820
when the regular power supply fails

133
00:06:05,820 --> 00:06:10,820
or importantly when voltage
drops to an unacceptable level.

134
00:06:11,160 --> 00:06:12,960
But the battery is finite, right?

135
00:06:12,960 --> 00:06:15,903
It has at some point the
battery will run out.

136
00:06:17,280 --> 00:06:20,280
A generator is a standby
secondary limited source

137
00:06:20,280 --> 00:06:23,220
of electrical power when
the power grid is down

138
00:06:23,220 --> 00:06:24,930
or inaccessible.

139
00:06:24,930 --> 00:06:27,000
We talked about this
way back when we talked

140
00:06:27,000 --> 00:06:29,190
about power and physical security, right?

141
00:06:29,190 --> 00:06:31,260
We said that the important
thing about a generator

142
00:06:31,260 --> 00:06:33,000
is we have to make sure
that fuel's available

143
00:06:33,000 --> 00:06:36,180
or generator's gonna run
on gasoline or on diesel,

144
00:06:36,180 --> 00:06:39,210
need to make sure that the
fuel is available to us.

145
00:06:39,210 --> 00:06:40,770
And then supplier diversity

146
00:06:40,770 --> 00:06:42,810
is the idea of having
more than one supplier

147
00:06:42,810 --> 00:06:45,513
or access to multiple power grids.

148
00:06:48,210 --> 00:06:51,570
Well, what about our transmission
in our wide area networks

149
00:06:51,570 --> 00:06:53,550
or transmission to the internet?

150
00:06:53,550 --> 00:06:56,610
Can we get resiliency in
our transmission routing?

151
00:06:56,610 --> 00:06:58,560
And the answer of course is gonna be, yes.

152
00:06:58,560 --> 00:06:59,760
Three options here.

153
00:06:59,760 --> 00:07:01,860
Alternate routing, diverse routing,

154
00:07:01,860 --> 00:07:04,830
and last mile circuit protection.

155
00:07:04,830 --> 00:07:07,380
In alternate routing,
we have multiple paths

156
00:07:07,380 --> 00:07:10,110
for our data to travel between two points.

157
00:07:10,110 --> 00:07:11,790
Now, the network can automatically

158
00:07:11,790 --> 00:07:14,460
reroute traffic to an alternate path,

159
00:07:14,460 --> 00:07:18,033
if the primary path becomes
unavailable or congested.

160
00:07:19,020 --> 00:07:21,480
In diverse routing, we
have data is transmitted

161
00:07:21,480 --> 00:07:25,380
over multiple geographically
diverse paths or routes.

162
00:07:25,380 --> 00:07:28,170
So taking different ways to get somewhere.

163
00:07:28,170 --> 00:07:30,510
And there are last mile circuit protection

164
00:07:30,510 --> 00:07:32,820
is where we have redundant
last mile circuit

165
00:07:32,820 --> 00:07:36,150
such as multiple fiber
optic or copper cables

166
00:07:36,150 --> 00:07:39,900
to provide backup paths
for data transmission

167
00:07:39,900 --> 00:07:43,530
in case of failure or outage
on one of our primary circuits.

168
00:07:43,530 --> 00:07:45,660
So alternate routing, diverse routing,

169
00:07:45,660 --> 00:07:47,733
and last mile circuit protection.

170
00:07:50,218 --> 00:07:52,290
Well, what if the location

171
00:07:52,290 --> 00:07:55,860
that we work out of isn't accessible

172
00:07:55,860 --> 00:07:57,450
or isn't viable anymore?

173
00:07:57,450 --> 00:08:01,260
Maybe there's been damage
to the building of some kind

174
00:08:01,260 --> 00:08:04,320
or just maybe there's a
been a natural disaster

175
00:08:04,320 --> 00:08:06,390
and we can't get to the building.

176
00:08:06,390 --> 00:08:08,820
So let's talk about
alternate physical sites

177
00:08:08,820 --> 00:08:11,160
physical sites or sites that are managed,

178
00:08:11,160 --> 00:08:14,343
maintained and monitored, you
know, by the organization.

179
00:08:15,480 --> 00:08:17,250
So some options we have for sites

180
00:08:17,250 --> 00:08:19,860
and this can be sites where
our users are working.

181
00:08:19,860 --> 00:08:22,410
It could be sites where we're
moving our data center to

182
00:08:22,410 --> 00:08:25,140
but these are alternate locations.

183
00:08:25,140 --> 00:08:28,830
AOC site is a site that
just has basic HVAC

184
00:08:28,830 --> 00:08:30,960
infrastructure that stands for heating,

185
00:08:30,960 --> 00:08:32,880
ventilation and air conditioning.

186
00:08:32,880 --> 00:08:34,260
So it's just a space, right?

187
00:08:34,260 --> 00:08:38,550
It has basic HVAC, but it's
really just an empty shell.

188
00:08:38,550 --> 00:08:41,310
There's no server related
or communications equipment.

189
00:08:41,310 --> 00:08:43,290
So if I have to move
there, I need everything.

190
00:08:43,290 --> 00:08:45,120
I need all my equipment,

191
00:08:45,120 --> 00:08:50,120
I need my cabling, I need
desks, I need computers,

192
00:08:50,160 --> 00:08:53,343
I need pencils, I need
paper, I need everything.

193
00:08:54,990 --> 00:08:57,900
A warm site is a site that does have HVAC,

194
00:08:57,900 --> 00:08:59,640
heating ventilation, air conditioning,

195
00:08:59,640 --> 00:09:00,900
and we've set it up a bit.

196
00:09:00,900 --> 00:09:03,703
It has servers and it has
communication infrastructure

197
00:09:03,703 --> 00:09:05,280
and equipment.

198
00:09:05,280 --> 00:09:07,710
Now, the systems probably
will need to be configured,

199
00:09:07,710 --> 00:09:09,600
updated they might need to be patched,

200
00:09:09,600 --> 00:09:11,250
they might need new antivirus

201
00:09:11,250 --> 00:09:13,503
and data will need to be restored.

202
00:09:15,000 --> 00:09:16,080
In a hot site,

203
00:09:16,080 --> 00:09:18,570
well, that's gonna have
HVAC, it'll have servers,

204
00:09:18,570 --> 00:09:20,430
it'll have our
communications infrastructure

205
00:09:20,430 --> 00:09:21,990
and it'll have equipment,

206
00:09:21,990 --> 00:09:24,270
and it's fully configured
and ready to operate.

207
00:09:24,270 --> 00:09:26,850
So we keep those systems
up to date, right?

208
00:09:26,850 --> 00:09:31,170
We're patching them if
we're making sure that

209
00:09:31,170 --> 00:09:33,090
you know, our antivirus is up to date.

210
00:09:33,090 --> 00:09:34,620
If there's anything that we've changed

211
00:09:34,620 --> 00:09:36,450
in our configuration management

212
00:09:36,450 --> 00:09:38,700
we're updating the systems as well.

213
00:09:38,700 --> 00:09:41,370
So the systems are fully
configured and ready to go.

214
00:09:41,370 --> 00:09:43,020
And data's been replicated.

215
00:09:43,020 --> 00:09:46,020
So either we've had asynchronous
or synchronous replication

216
00:09:46,020 --> 00:09:49,080
so we should be back in
business pretty quick.

217
00:09:49,080 --> 00:09:50,970
And lastly is a mirrored site.

218
00:09:50,970 --> 00:09:52,830
And a mirrored site isn't identical

219
00:09:52,830 --> 00:09:56,220
or nearly identical site that's
already operational, right?

220
00:09:56,220 --> 00:09:59,490
It's operational in concert
with the primary site

221
00:09:59,490 --> 00:10:01,110
on a load balancing basis.

222
00:10:01,110 --> 00:10:03,540
So we lose one site,
we're still operational

223
00:10:03,540 --> 00:10:04,533
at the other site.

224
00:10:07,170 --> 00:10:08,550
Then we have a couple of options

225
00:10:08,550 --> 00:10:11,490
for alternate third party
sites, a mobile site,

226
00:10:11,490 --> 00:10:14,943
a reciprocal site, and a cloud-based site.

227
00:10:16,260 --> 00:10:19,680
A mobile site is going to be
a transportable module unit

228
00:10:19,680 --> 00:10:21,660
that will show up at your doorstep, right?

229
00:10:21,660 --> 00:10:24,030
And it has your pre-ordered
hardware and software

230
00:10:24,030 --> 00:10:26,130
and you've probably been
sending your backups

231
00:10:26,130 --> 00:10:29,130
to whoever owns that mobile site.

232
00:10:29,130 --> 00:10:32,010
That sounds great, except
there's a lot of situations

233
00:10:32,010 --> 00:10:34,170
where it's really not gonna work

234
00:10:34,170 --> 00:10:37,230
because the delivery site
must provide access roads,

235
00:10:37,230 --> 00:10:41,760
water, waste disposal,
power and connectivity.

236
00:10:41,760 --> 00:10:44,910
So particularly in a natural disaster,

237
00:10:44,910 --> 00:10:46,623
it might not be your best choice.

238
00:10:47,580 --> 00:10:49,860
A reciprocal site is based on an agreement

239
00:10:49,860 --> 00:10:54,003
to have access to or use another
organization's facilities.

240
00:10:55,020 --> 00:10:56,010
That's a little scary.

241
00:10:56,010 --> 00:10:58,890
But you know, often used
by very small companies

242
00:10:58,890 --> 00:11:00,870
but you know, you don't necessarily know

243
00:11:00,870 --> 00:11:03,000
what they're gonna have available to you.

244
00:11:03,000 --> 00:11:04,560
You may not be confident

245
00:11:04,560 --> 00:11:08,100
in their privacy or security controls.

246
00:11:08,100 --> 00:11:09,660
You know, their physical controls,

247
00:11:09,660 --> 00:11:12,600
their environmental controls,
and just the idea that,

248
00:11:12,600 --> 00:11:15,810
okay, if I go down, I'm gonna
use somebody else's facility.

249
00:11:15,810 --> 00:11:17,460
That seems to me to make much more sense,

250
00:11:17,460 --> 00:11:19,920
if it's just I need office
space and desk space

251
00:11:19,920 --> 00:11:24,900
opposed to I'm gonna put my
servers in this new location.

252
00:11:24,900 --> 00:11:27,540
And lastly, we have DRAAS,

253
00:11:27,540 --> 00:11:30,090
which is disaster recovery as a service.

254
00:11:30,090 --> 00:11:32,430
Now cloud-based disaster
recovery as a service

255
00:11:32,430 --> 00:11:36,600
offers full recovery in a
cloud-based environment.

256
00:11:36,600 --> 00:11:40,020
So lots of options for resiliency, right?

257
00:11:40,020 --> 00:11:44,247
Site, in power, in system, in data, right?

258
00:11:44,247 --> 00:11:47,400
And we always have to keep
in mind that systems fail

259
00:11:47,400 --> 00:11:49,080
and problems occur

260
00:11:49,080 --> 00:11:51,510
and we have to be ready
to continue to operate.

261
00:11:51,510 --> 00:11:53,520
And that's really the goal of resiliency

262
00:11:53,520 --> 00:11:55,200
to be able to continue to operate

263
00:11:55,200 --> 00:11:57,543
in abnormal operating conditions.

264
00:11:58,440 --> 00:12:01,170
That my friends, brings us
to a three second challenge.

265
00:12:01,170 --> 00:12:03,930
Five challenge questions,
three seconds each.

266
00:12:03,930 --> 00:12:04,763
No question.

267
00:12:04,763 --> 00:12:06,090
You know how to do it.

268
00:12:06,090 --> 00:12:09,780
Question one, the capability
to continue to operate

269
00:12:09,780 --> 00:12:11,940
in abnormal conditions.

270
00:12:11,940 --> 00:12:16,590
One, two, three, that's
gonna be resiliency.

271
00:12:16,590 --> 00:12:20,220
Number two, this technology commonly used

272
00:12:20,220 --> 00:12:21,903
for fault tolerance.

273
00:12:23,490 --> 00:12:27,840
Showed you two versions One, two, three.

274
00:12:27,840 --> 00:12:29,103
And that's gonna be RAID.

275
00:12:30,480 --> 00:12:32,520
Number three, when data is transmitted

276
00:12:32,520 --> 00:12:36,280
over multiple geographically
diverse paths or routes

277
00:12:38,340 --> 00:12:43,113
One, two, three, it's
gonna be diverse routing.

278
00:12:44,640 --> 00:12:47,370
Number four, this site has HVAC

279
00:12:47,370 --> 00:12:50,790
servers and communications
infrastructure and equipment.

280
00:12:50,790 --> 00:12:52,920
Systems might need to be configured.

281
00:12:52,920 --> 00:12:57,920
Data needs to be
restored. One, two, three,

282
00:12:58,440 --> 00:13:01,410
That's it gonna be, it's a warm site.

283
00:13:01,410 --> 00:13:03,900
And lastly, number five, the process

284
00:13:03,900 --> 00:13:07,950
of writing data on two
partitions on separate discs.

285
00:13:07,950 --> 00:13:09,873
One, two, three.

286
00:13:10,890 --> 00:13:13,440
And that's known as disc mirroring.

287
00:13:13,440 --> 00:13:15,030
And if we add another controller,

288
00:13:15,030 --> 00:13:16,653
it would be disc duplexing.

289
00:13:18,270 --> 00:13:20,070
Okay, let's do a security in action.

290
00:13:20,070 --> 00:13:25,070
This one is about uptime, 99.999%.

291
00:13:25,290 --> 00:13:27,540
Your organization is
negotiating a contract

292
00:13:27,540 --> 00:13:28,860
with a new client.

293
00:13:28,860 --> 00:13:31,470
They're requiring the
service that you're proposing

294
00:13:31,470 --> 00:13:34,677
to have an uptime of 99.999%.

295
00:13:37,530 --> 00:13:38,970
Now, before agreeing to these terms,

296
00:13:38,970 --> 00:13:41,370
you recommend an internal discussion

297
00:13:41,370 --> 00:13:44,100
to evaluate our current capacity

298
00:13:44,100 --> 00:13:46,200
and to determine the investment required

299
00:13:46,200 --> 00:13:49,800
to guarantee this level of availability.

300
00:13:49,800 --> 00:13:52,440
Now, the standard configuration
what you have right now

301
00:13:52,440 --> 00:13:56,400
is RAID 5 with the ability
to hot swap drives.

302
00:13:56,400 --> 00:13:57,270
So if a drive fails

303
00:13:57,270 --> 00:13:59,370
you can pull one out and put the other in.

304
00:14:00,750 --> 00:14:01,860
But is that enough?

305
00:14:01,860 --> 00:14:05,070
Is that gonna get you to 99.999

306
00:14:05,070 --> 00:14:06,750
plus are there other things you need?

307
00:14:06,750 --> 00:14:08,160
So my question to you

308
00:14:08,160 --> 00:14:11,253
is what discussion topics would you raise?

309
00:14:12,390 --> 00:14:13,380
So a quick recap,

310
00:14:13,380 --> 00:14:16,590
you are negotiating a
contract with a new client.

311
00:14:16,590 --> 00:14:21,420
They're saying, "I want you
to be up 99.999% of the time."

312
00:14:21,420 --> 00:14:24,660
And you're saying, "Okay,
well, but before we do that

313
00:14:24,660 --> 00:14:27,780
we really need to evaluate
our current capacity

314
00:14:27,780 --> 00:14:30,690
and see what kind of
investment would be required

315
00:14:30,690 --> 00:14:33,510
to guarantee this level of availability."

316
00:14:33,510 --> 00:14:36,210
Maybe this contract is not worth it.

317
00:14:36,210 --> 00:14:37,200
What do you do right now?

318
00:14:37,200 --> 00:14:39,840
Well, your current
configuration is RAID 5.

319
00:14:39,840 --> 00:14:42,510
That's gonna be disc striping with parody

320
00:14:42,510 --> 00:14:44,940
with the ability to hot swap a drive.

321
00:14:44,940 --> 00:14:47,160
Hot swap means you can just
take a drive out on the fly

322
00:14:47,160 --> 00:14:48,540
and put a new one in.

323
00:14:48,540 --> 00:14:50,220
So what topics are you gonna raise?

324
00:14:50,220 --> 00:14:51,690
Go ahead and put me on pause,

325
00:14:51,690 --> 00:14:54,390
write down those discussion
topics, then come on back.

326
00:14:57,090 --> 00:14:59,910
Well, 99.999% of time,

327
00:14:59,910 --> 00:15:01,830
and you see that in
marketing all the time,

328
00:15:01,830 --> 00:15:04,770
implies no more than 5.26 minutes

329
00:15:04,770 --> 00:15:06,900
of unavailability per year.

330
00:15:06,900 --> 00:15:09,963
That's a really tough nut to achieve.

331
00:15:11,760 --> 00:15:14,640
Now, RAID is a disk-space
fault tolerance strategy.

332
00:15:14,640 --> 00:15:15,750
Pretty good strategy, right?

333
00:15:15,750 --> 00:15:17,520
Especially with our hot swap drives.

334
00:15:17,520 --> 00:15:20,130
But it doesn't extend to
any other components, right?

335
00:15:20,130 --> 00:15:22,323
It's only about drives.

336
00:15:23,220 --> 00:15:26,220
So investments in high
availability systems

337
00:15:26,220 --> 00:15:27,720
are gonna need to be made

338
00:15:27,720 --> 00:15:30,870
either asymmetric, high
availability or symmetric.

339
00:15:30,870 --> 00:15:33,180
Remember, we can either
have an active passive pair

340
00:15:33,180 --> 00:15:35,013
or an active active pair.

341
00:15:36,570 --> 00:15:40,140
The consideration should also
be given to site resiliency.

342
00:15:40,140 --> 00:15:42,490
What happens if there's
a problem at your site?

343
00:15:43,770 --> 00:15:46,680
And then additional downtime
considerations include,

344
00:15:46,680 --> 00:15:48,893
necessary maintenance windows for upgrades

345
00:15:48,893 --> 00:15:51,660
and for patch management.

346
00:15:51,660 --> 00:15:54,360
So trying to achieve 5.26 minutes

347
00:15:54,360 --> 00:15:58,410
of unavailability per year
really is a very high bar.

348
00:15:58,410 --> 00:16:00,780
So before you're agreeing
to that contract,

349
00:16:00,780 --> 00:16:02,160
right you really wanna go through

350
00:16:02,160 --> 00:16:03,510
what are all the things that you need

351
00:16:03,510 --> 00:16:04,860
how is it different than what you have?

352
00:16:04,860 --> 00:16:06,030
So a gap analysis

353
00:16:06,030 --> 00:16:08,880
and what type of investment
is gonna be required.

354
00:16:08,880 --> 00:16:12,150
And then ultimately doing
a cost benefit analysis

355
00:16:12,150 --> 00:16:13,170
doing all that.

356
00:16:13,170 --> 00:16:15,690
That's definitely security and action.

357
00:16:15,690 --> 00:16:16,830
It's a pretty big word cloud.

358
00:16:16,830 --> 00:16:18,780
Once again, take your time,

359
00:16:18,780 --> 00:16:23,190
make sure that you understand
all of these concepts

360
00:16:23,190 --> 00:16:24,023
and terms.

361
00:16:24,023 --> 00:16:25,890
You can speak to them confidently.

362
00:16:25,890 --> 00:16:28,380
And when you're ready, come
on over to the next lesson.

363
00:16:28,380 --> 00:16:30,180
I'll be waiting for you right there.
