1
00:00:00,000 --> 00:00:00,900
In this lesson,

2
00:00:00,900 --> 00:00:02,940
we're going to cover high availability.

3
00:00:02,940 --> 00:00:04,920
These days, our organizations expect

4
00:00:04,920 --> 00:00:07,830
to have our services up and running pretty much non-stop.

5
00:00:07,830 --> 00:00:09,150
And the only way to achieve that

6
00:00:09,150 --> 00:00:10,620
is by creating an architecture

7
00:00:10,620 --> 00:00:13,080
that will support high levels of availability.

8
00:00:13,080 --> 00:00:15,690
Now, high availability refers to the ability of a service

9
00:00:15,690 --> 00:00:18,360
to be continuously available by minimizing the downtime

10
00:00:18,360 --> 00:00:20,070
to the lowest amount possible.

11
00:00:20,070 --> 00:00:21,570
To achieve high availability,

12
00:00:21,570 --> 00:00:23,640
our systems should be designed using load balancing

13
00:00:23,640 --> 00:00:26,280
or clustering, have redundancy built into the system,

14
00:00:26,280 --> 00:00:28,770
and if we're operating in a cloud-based environment,

15
00:00:28,770 --> 00:00:31,290
we should also consider using a multi-cloud environment

16
00:00:31,290 --> 00:00:32,700
to achieve an operational environment

17
00:00:32,700 --> 00:00:34,590
that can withstand multiple points of failure

18
00:00:34,590 --> 00:00:36,900
before our service becomes unavailable.

19
00:00:36,900 --> 00:00:38,670
Now, whenever we talk about availability,

20
00:00:38,670 --> 00:00:40,680
it's important to remember that we measure our levels

21
00:00:40,680 --> 00:00:44,010
of availability in terms of something known as uptime.

22
00:00:44,010 --> 00:00:45,810
Uptime is the number of minutes or hours

23
00:00:45,810 --> 00:00:48,570
that your system remains online over a given period,

24
00:00:48,570 --> 00:00:50,640
and this uptime is usually going to be expressed

25
00:00:50,640 --> 00:00:51,960
as a percentage.

26
00:00:51,960 --> 00:00:54,060
The gold standard in availability is set

27
00:00:54,060 --> 00:00:55,170
by achieving something known

28
00:00:55,170 --> 00:00:57,390
as the five nines of availability.

29
00:00:57,390 --> 00:00:59,310
The five nines of availability refers

30
00:00:59,310 --> 00:01:02,580
to achieving a 99.999% uptime,

31
00:01:02,580 --> 00:01:04,830
which means that your services can only have a maximum

32
00:01:04,830 --> 00:01:07,110
of about five minutes of downtime per year,

33
00:01:07,110 --> 00:01:08,790
which as you can probably imagine,

34
00:01:08,790 --> 00:01:10,800
isn't a whole lot of downtime.

35
00:01:10,800 --> 00:01:12,390
Now, in some cloud-based networks,

36
00:01:12,390 --> 00:01:14,250
organizations have gone even further

37
00:01:14,250 --> 00:01:16,860
by setting their availability target at six nines

38
00:01:16,860 --> 00:01:20,760
of availability, or 99.9999%,

39
00:01:20,760 --> 00:01:23,400
which equates to only 31 seconds of downtime

40
00:01:23,400 --> 00:01:25,050
each and every year.

41
00:01:25,050 --> 00:01:26,160
Now, as you can imagine,

42
00:01:26,160 --> 00:01:27,810
most organizations are going to need

43
00:01:27,810 --> 00:01:30,720
to take their systems offline for more than 31 seconds

44
00:01:30,720 --> 00:01:33,780
per year and probably more than five minutes each year too

45
00:01:33,780 --> 00:01:36,240
in order to install security patches on their servers

46
00:01:36,240 --> 00:01:37,890
or to replace a failed hard drive

47
00:01:37,890 --> 00:01:40,320
or to install a new router or switch into the network

48
00:01:40,320 --> 00:01:42,060
when the old one has failed.

49
00:01:42,060 --> 00:01:44,790
So how can we maintain a high level of availability

50
00:01:44,790 --> 00:01:47,070
while still being able to take down some of our systems

51
00:01:47,070 --> 00:01:49,260
to perform our maintenance and repairs?

52
00:01:49,260 --> 00:01:51,000
Well, the first way we can do this in order

53
00:01:51,000 --> 00:01:52,860
to achieve higher levels of availability

54
00:01:52,860 --> 00:01:54,687
is to prevent our systems from becoming overloaded

55
00:01:54,687 --> 00:01:56,820
by using load balancing.

56
00:01:56,820 --> 00:01:59,190
Load balancing is the process of distributing workloads

57
00:01:59,190 --> 00:02:01,140
across multiple computing resources

58
00:02:01,140 --> 00:02:04,230
to optimize the resource use, maximize their throughput,

59
00:02:04,230 --> 00:02:05,670
minimize their response time,

60
00:02:05,670 --> 00:02:08,520
and prevent the overloading of any single resource.

61
00:02:08,520 --> 00:02:11,550
Load balancers achieve this by using complex algorithms

62
00:02:11,550 --> 00:02:13,470
to distribute incoming requests to servers

63
00:02:13,470 --> 00:02:14,850
that are capable of fulfilling them

64
00:02:14,850 --> 00:02:16,770
efficiently and effectively.

65
00:02:16,770 --> 00:02:18,990
For example, if I'm running a small website

66
00:02:18,990 --> 00:02:21,270
to host my blog with only a few dozen readers,

67
00:02:21,270 --> 00:02:23,790
I can probably get by with just having a single server

68
00:02:23,790 --> 00:02:25,980
that can handle that amount of simultaneous requests

69
00:02:25,980 --> 00:02:27,000
that I'm expecting to receive

70
00:02:27,000 --> 00:02:29,160
without causing my server to crash.

71
00:02:29,160 --> 00:02:31,530
But if my blog starts to gain popularity

72
00:02:31,530 --> 00:02:33,240
and now I have a few thousands people who are trying

73
00:02:33,240 --> 00:02:35,790
to read it, a single server is going to be insufficient

74
00:02:35,790 --> 00:02:37,680
to handle all of that additional load.

75
00:02:37,680 --> 00:02:40,470
So instead, I can set up two or three servers

76
00:02:40,470 --> 00:02:42,720
that will all take turns responding to my visitors,

77
00:02:42,720 --> 00:02:45,840
to be able to ensure that no single server gets too overloaded.

78
00:02:45,840 --> 00:02:47,970
This is what load balancing allows us to do,

79
00:02:47,970 --> 00:02:50,580
because all the requests go to the load balancer first,

80
00:02:50,580 --> 00:02:53,400
and then the load balancer will redirect that user's request

81
00:02:53,400 --> 00:02:55,620
to one of my three servers in order to respond

82
00:02:55,620 --> 00:02:57,780
to the request in a more timely manner.

83
00:02:57,780 --> 00:02:59,790
Now, second, we have clustering.

84
00:02:59,790 --> 00:03:01,560
Clustering can be set up to help your systems

85
00:03:01,560 --> 00:03:04,020
to handle more traffic at the same time.

86
00:03:04,020 --> 00:03:06,240
Clustering refers to the use of multiple computers,

87
00:03:06,240 --> 00:03:09,030
multiple storage devices, or redundant network connections

88
00:03:09,030 --> 00:03:11,250
that all work together as a single system

89
00:03:11,250 --> 00:03:13,020
to provide higher levels of availability,

90
00:03:13,020 --> 00:03:16,560
reliability, and scalability for your enterprise systems.

91
00:03:16,560 --> 00:03:19,140
Clustering is more about keeping an application available

92
00:03:19,140 --> 00:03:21,000
even in the event of a hardware failure

93
00:03:21,000 --> 00:03:21,833
in order to ensure

94
00:03:21,833 --> 00:03:24,000
that there is no single points of failure.

95
00:03:24,000 --> 00:03:26,550
While load balancing can help handle all the excess traffic

96
00:03:26,550 --> 00:03:29,460
by distributing it, clustering will provide redundancy

97
00:03:29,460 --> 00:03:31,110
in the event of a system failure

98
00:03:31,110 --> 00:03:32,910
to ensure that the continuity of service is going

99
00:03:32,910 --> 00:03:34,200
to be maintained.

100
00:03:34,200 --> 00:03:36,780
Load balancing and clustering can also be combined together

101
00:03:36,780 --> 00:03:38,160
in the same architecture

102
00:03:38,160 --> 00:03:39,750
to provide an even more robust system

103
00:03:39,750 --> 00:03:42,150
for maintaining higher levels of availability.

104
00:03:42,150 --> 00:03:43,980
Under this type of combined design,

105
00:03:43,980 --> 00:03:46,590
load balancing will manage and optimize your resources

106
00:03:46,590 --> 00:03:47,880
under normal conditions,

107
00:03:47,880 --> 00:03:49,470
but then clustering can take over

108
00:03:49,470 --> 00:03:51,690
to ensure the service remains available in the event

109
00:03:51,690 --> 00:03:54,390
of a component failure or a system failure.

110
00:03:54,390 --> 00:03:56,910
Third, our high availability systems should be built

111
00:03:56,910 --> 00:03:58,530
with redundancy in mind.

112
00:03:58,530 --> 00:04:01,050
Now, redundancy is the duplication of critical components

113
00:04:01,050 --> 00:04:02,460
or functions of a system

114
00:04:02,460 --> 00:04:04,560
with the intention of increasing the reliability

115
00:04:04,560 --> 00:04:05,970
of a given system.

116
00:04:05,970 --> 00:04:07,920
To create a highly available infrastructure,

117
00:04:07,920 --> 00:04:09,810
redundancy has to be built into your designs

118
00:04:09,810 --> 00:04:12,270
by installing or adding multiple power supplies,

119
00:04:12,270 --> 00:04:14,850
network connections, servers, software services,

120
00:04:14,850 --> 00:04:17,760
or service providers to your architecture's design.

121
00:04:17,760 --> 00:04:19,500
By using redundant power supplies,

122
00:04:19,500 --> 00:04:21,600
you can ensure that the failure of one power source

123
00:04:21,600 --> 00:04:24,480
is not going to affect the continuity of your other services.

124
00:04:24,480 --> 00:04:26,490
Now, this type of power supply redundancy

125
00:04:26,490 --> 00:04:28,590
can be achieved by installing two power supplies

126
00:04:28,590 --> 00:04:30,270
inside of a server directly,

127
00:04:30,270 --> 00:04:32,550
or it can be achieved by ensuring your facility

128
00:04:32,550 --> 00:04:34,110
has multiple power sources

129
00:04:34,110 --> 00:04:36,870
by using an uninterrupted power supply, or UPS,

130
00:04:36,870 --> 00:04:39,210
a backup generator, or maintaining connections

131
00:04:39,210 --> 00:04:41,520
to two or more power grids simultaneously

132
00:04:41,520 --> 00:04:43,140
to ensure your systems always have access

133
00:04:43,140 --> 00:04:45,090
to the power they need to operate.

134
00:04:45,090 --> 00:04:47,010
In addition to redundant power supplies,

135
00:04:47,010 --> 00:04:47,843
it's also important

136
00:04:47,843 --> 00:04:50,280
that you maintain multiple network connections or pathways

137
00:04:50,280 --> 00:04:51,960
by using multiple cabled connections

138
00:04:51,960 --> 00:04:54,630
or a cable connection and a wireless connection

139
00:04:54,630 --> 00:04:56,550
in order to prevent any single points of failure

140
00:04:56,550 --> 00:04:59,670
from disrupting your service to your network's connectivity.

141
00:04:59,670 --> 00:05:01,920
Now, another form of redundancy comes in the form

142
00:05:01,920 --> 00:05:03,360
of servers and services,

143
00:05:03,360 --> 00:05:05,370
which can be configured to operate in a load balanced

144
00:05:05,370 --> 00:05:07,766
or clustered architecture to help prevent downtime

145
00:05:07,766 --> 00:05:10,050
by ensuring that a backup will always be available

146
00:05:10,050 --> 00:05:11,520
in case one of your servers suffers

147
00:05:11,520 --> 00:05:13,230
from a catastrophic failure.

148
00:05:13,230 --> 00:05:15,690
Similarly, deploying redundant software services

149
00:05:15,690 --> 00:05:17,430
or components can be used to ensure

150
00:05:17,430 --> 00:05:18,990
that if one instance fails,

151
00:05:18,990 --> 00:05:20,700
then the remaining workload can be shifted

152
00:05:20,700 --> 00:05:22,680
to another instance seamlessly.

153
00:05:22,680 --> 00:05:24,330
Now, additionally, if you're reliant

154
00:05:24,330 --> 00:05:25,590
on external service providers

155
00:05:25,590 --> 00:05:27,570
to provide a critically important service to you,

156
00:05:27,570 --> 00:05:29,490
like your internet service provider does,

157
00:05:29,490 --> 00:05:31,380
then you should probably protect your organization

158
00:05:31,380 --> 00:05:34,350
from an outage by using two or more service providers

159
00:05:34,350 --> 00:05:36,810
so that you always have a readily available backup.

160
00:05:36,810 --> 00:05:38,580
This kind of provider diversity means

161
00:05:38,580 --> 00:05:40,920
that if one provider experiences an internal issue,

162
00:05:40,920 --> 00:05:43,380
the organization's operations will be able to continue

163
00:05:43,380 --> 00:05:45,570
by using another provider's services.

164
00:05:45,570 --> 00:05:47,640
For example, at DionTraining.com,

165
00:05:47,640 --> 00:05:50,010
we use a credit card service provider known as Stripe

166
00:05:50,010 --> 00:05:51,510
to collect payments from our students

167
00:05:51,510 --> 00:05:53,940
when they buy their exam vouchers on our website.

168
00:05:53,940 --> 00:05:56,520
But we also have a secondary provider already configured

169
00:05:56,520 --> 00:05:58,890
and ready for us to enable if Stripe has some kind

170
00:05:58,890 --> 00:06:00,540
of a catastrophic outage.

171
00:06:00,540 --> 00:06:02,490
Normally, we prefer to rely on Stripe

172
00:06:02,490 --> 00:06:04,260
because they provide a great level of service

173
00:06:04,260 --> 00:06:06,510
at a lower price than our secondary provider.

174
00:06:06,510 --> 00:06:08,820
But if Stripe goes offline for some reason,

175
00:06:08,820 --> 00:06:09,750
we're already configured

176
00:06:09,750 --> 00:06:12,240
that our backup credit card processor can take over

177
00:06:12,240 --> 00:06:14,730
until our Stripe-based service is once again ready

178
00:06:14,730 --> 00:06:17,340
to start conducting normal operations again.

179
00:06:17,340 --> 00:06:19,650
Another example might be having two domain controllers

180
00:06:19,650 --> 00:06:21,570
in your enterprise network environment.

181
00:06:21,570 --> 00:06:23,040
Under this kind of a design,

182
00:06:23,040 --> 00:06:24,390
you can set up your domain controllers

183
00:06:24,390 --> 00:06:26,550
as a primary and secondary controller.

184
00:06:26,550 --> 00:06:29,040
And if you need to restart one to apply a security patch,

185
00:06:29,040 --> 00:06:31,350
the other controller can continue to maintain the services

186
00:06:31,350 --> 00:06:33,030
for all of your end users.

187
00:06:33,030 --> 00:06:34,950
Having redundant hardware for everything

188
00:06:34,950 --> 00:06:36,540
would become really expensive,

189
00:06:36,540 --> 00:06:38,220
because effectively, we are doubling the cost

190
00:06:38,220 --> 00:06:39,360
to build out our network.

191
00:06:39,360 --> 00:06:41,970
So it's important that each design decision is considered

192
00:06:41,970 --> 00:06:43,710
to decide if redundancy is required

193
00:06:43,710 --> 00:06:45,240
for a particular piece of hardware

194
00:06:45,240 --> 00:06:47,370
or if you could build the desired level of redundancy

195
00:06:47,370 --> 00:06:50,280
by using software or cloud-based services instead.

196
00:06:50,280 --> 00:06:52,230
Either way, though, by adding additional layers

197
00:06:52,230 --> 00:06:54,360
of redundancy to your overall architecture,

198
00:06:54,360 --> 00:06:55,650
you're going to be able to help minimize

199
00:06:55,650 --> 00:06:57,750
the potential points of failures in your systems

200
00:06:57,750 --> 00:07:00,390
and ensure that your system or services remain available

201
00:07:00,390 --> 00:07:02,580
despite any unforeseen issues.

202
00:07:02,580 --> 00:07:04,890
Fourth, we can utilize multi-cloud systems

203
00:07:04,890 --> 00:07:07,470
to distribute our data, applications, and services

204
00:07:07,470 --> 00:07:09,570
across several different cloud-based environments

205
00:07:09,570 --> 00:07:11,610
as an additional form of redundancy.

206
00:07:11,610 --> 00:07:13,740
By implementing a multi-cloud architecture,

207
00:07:13,740 --> 00:07:15,960
you can mitigate the risk of a single point of failure,

208
00:07:15,960 --> 00:07:17,430
because if one of your cloud providers

209
00:07:17,430 --> 00:07:18,600
experiences an outage,

210
00:07:18,600 --> 00:07:20,040
then your organization's workload

211
00:07:20,040 --> 00:07:22,650
can be rapidly transitioned to another cloud provider

212
00:07:22,650 --> 00:07:25,260
with minimal disruptions to your operations.

213
00:07:25,260 --> 00:07:26,730
Using a multi-cloud system

214
00:07:26,730 --> 00:07:28,710
will also provide you with some additional flexibility

215
00:07:28,710 --> 00:07:31,650
for scaling your operations and for optimizing your costs,

216
00:07:31,650 --> 00:07:33,030
because different service providers

217
00:07:33,030 --> 00:07:36,060
may offer different pricing structures for similar services.

218
00:07:36,060 --> 00:07:38,220
So you can opt to use the least expensive one

219
00:07:38,220 --> 00:07:40,200
that meets your needs as your primary provider

220
00:07:40,200 --> 00:07:42,720
for a given service while maintaining the other provider

221
00:07:42,720 --> 00:07:46,320
as your secondary or backup provider for that same service.

222
00:07:46,320 --> 00:07:49,110
A multi-cloud design also helps your organization avoid

223
00:07:49,110 --> 00:07:51,210
any kind of potential vendor lock-in issues,

224
00:07:51,210 --> 00:07:53,700
because it provides you with more options and leverage

225
00:07:53,700 --> 00:07:56,160
when it comes time to negotiate your service terms

226
00:07:56,160 --> 00:07:57,840
or if you need to migrate your services

227
00:07:57,840 --> 00:08:00,000
to another cloud provider due to an outage

228
00:08:00,000 --> 00:08:01,710
or other types of issues.

229
00:08:01,710 --> 00:08:04,350
However, while implementing a multi-cloud strategy,

230
00:08:04,350 --> 00:08:05,190
it's going to be important

231
00:08:05,190 --> 00:08:06,990
that you ensure proper data management,

232
00:08:06,990 --> 00:08:08,100
unified threat management,

233
00:08:08,100 --> 00:08:10,170
and consistent policy enforcement is being used

234
00:08:10,170 --> 00:08:12,150
across all of your cloud-based environments

235
00:08:12,150 --> 00:08:14,610
so that your organization can maintain a higher security

236
00:08:14,610 --> 00:08:16,200
and compliance posture.

237
00:08:16,200 --> 00:08:18,780
So remember, to achieve high availability,

238
00:08:18,780 --> 00:08:20,670
you're going to need to conduct some strategic planning

239
00:08:20,670 --> 00:08:22,950
in order to design a more robust system architecture

240
00:08:22,950 --> 00:08:24,360
for your organization.

241
00:08:24,360 --> 00:08:26,250
By understanding and implementing load balancing

242
00:08:26,250 --> 00:08:28,860
and clustering, instituting redundancy and multiple layers

243
00:08:28,860 --> 00:08:30,300
of the operational framework,

244
00:08:30,300 --> 00:08:32,309
and utilizing a multi-cloud approach,

245
00:08:32,309 --> 00:08:34,559
organizations can significantly reduce their risk

246
00:08:34,559 --> 00:08:37,890
of service disruptions and the associated downtime costs.

247
00:08:37,890 --> 00:08:40,080
This kind of proactive approach helps to safeguard

248
00:08:40,080 --> 00:08:43,049
your organization's operational continuity and reliability

249
00:08:43,049 --> 00:08:45,540
while also fortifying its reputation and credibility

250
00:08:45,540 --> 00:08:48,290
in an increasingly competitive operational environment.

