1
1

00:00:00,400  -->  00:00:02,600
<v ->In the last lesson I showed you a couple of diagrams</v>
2

2

00:00:02,600  -->  00:00:04,150
of redundant networks,
3

3

00:00:04,150  -->  00:00:06,670
but one of the things we had to think about in this lesson
4

4

00:00:06,670  -->  00:00:08,160
are the considerations we have
5

5

00:00:08,160  -->  00:00:10,640
when we start designing these redundant networks.
6

6

00:00:10,640  -->  00:00:11,970
First you need to ask yourself,
7

7

00:00:11,970  -->  00:00:13,970
are you going to use redundancy in the network,
8

8

00:00:13,970  -->  00:00:16,670
and if so, where, and how?
9

9

00:00:16,670  -->  00:00:19,590
So are you going to do it from a module or a parts perspective?
10

10

00:00:19,590  -->  00:00:21,750
For instance, are you going to have multiple power supplies,
11

11

00:00:21,750  -->  00:00:25,450
multiple network interface devices, multiple hard drives,
12

12

00:00:25,450  -->  00:00:27,840
or are you going to look at it more from a chassis redundancy
13

13

00:00:27,840  -->  00:00:31,040
and have two sets of routers or two sets of switches?
14

14

00:00:31,040  -->  00:00:32,555
These are things you have to think about.
15

15

00:00:32,555  -->  00:00:34,620
Which one of these are you going to use,
16

16

00:00:34,620  -->  00:00:36,299
because each one is going to affect the cost
17

17

00:00:36,299  -->  00:00:38,960
of your network, based on the decisions you make.
18

18

00:00:38,960  -->  00:00:40,710
You have to be able to make a good business case
19

19

00:00:40,710  -->  00:00:42,870
for which one you're going to use, and why.
20

20

00:00:42,870  -->  00:00:43,960
For instance, if you could just have
21

21

00:00:43,960  -->  00:00:46,582
a second network interface card or a second power supply,
22

22

00:00:46,582  -->  00:00:48,430
that's going to be a lot cheaper
23

23

00:00:48,430  -->  00:00:49,810
than having to have an entire switch
24

24

00:00:49,810  -->  00:00:51,680
or an entire extra router there.
25

25

00:00:51,680  -->  00:00:53,420
Now, each of those switches and routers,
26

26

00:00:53,420  -->  00:00:56,240
some of these can cost 3 or 4 or $5,000,
27

27

00:00:56,240  -->  00:00:57,660
and so it might be a lot cheaper
28

28

00:00:57,660  -->  00:00:59,157
to have a redundant power supply, right,
29

29

00:00:59,157  -->  00:01:00,820
and so these are the things you have to think about
30

30

00:01:00,820  -->  00:01:03,040
and weigh as you're building your networks.
31

31

00:01:03,040  -->  00:01:04,101
Another thing you have to think about
32

32

00:01:04,101  -->  00:01:05,730
is software redundancy,
33

33

00:01:05,730  -->  00:01:07,960
and which features of those are going to be appropriate.
34

34

00:01:07,960  -->  00:01:10,590
Sometimes you can solve a lot of these redundancy problems
35

35

00:01:10,590  -->  00:01:13,090
by using software as opposed to hardware.
36

36

00:01:13,090  -->  00:01:15,660
For example, if you have a virtual network setup,
37

37

00:01:15,660  -->  00:01:17,120
you could just put in a virtual switch
38

38

00:01:17,120  -->  00:01:18,510
or a virtual router in there,
39

39

00:01:18,510  -->  00:01:19,480
and that way you don't have to bring
40

40

00:01:19,480  -->  00:01:21,600
another real router or real switch in,
41

41

00:01:21,600  -->  00:01:23,810
that can save you a lot of money.
42

42

00:01:23,810  -->  00:01:26,180
There's also a lot of other software solutions out there,
43

43

00:01:26,180  -->  00:01:27,500
like a software RAID,
44

44

00:01:27,500  -->  00:01:28,700
that will give you additional redundancy
45

45

00:01:28,700  -->  00:01:29,990
for your storage devices,
46

46

00:01:29,990  -->  00:01:32,700
as opposed to putting in an extra hard drive chassis,
47

47

00:01:32,700  -->  00:01:35,440
or another RAID array or storage area network.
48

48

00:01:35,440  -->  00:01:36,670
Also, these are the types of things
49

49

00:01:36,670  -->  00:01:37,503
you have to be thinking about
50

50

00:01:37,503  -->  00:01:39,570
as you're building out your network, right?
51

51

00:01:39,570  -->  00:01:41,010
When you think about your protocols,
52

52

00:01:41,010  -->  00:01:42,400
what protocol characteristics
53

53

00:01:42,400  -->  00:01:44,540
are going to affect your design requirements?
54

54

00:01:44,540  -->  00:01:46,650
This is really important if you're designing things,
55

55

00:01:46,650  -->  00:01:47,520
and you're using something like
56

56

00:01:47,520  -->  00:01:49,780
TCP versus UDP in your designs,
57

57

00:01:49,780  -->  00:01:52,720
because TCP has that additional redundancy
58

58

00:01:52,720  -->  00:01:55,360
by resending packets, where UDP doesn't,
59

59

00:01:55,360  -->  00:01:57,720
this is something you have to consider as well.
60

60

00:01:57,720  -->  00:01:59,580
As you design all these different things,
61

61

00:01:59,580  -->  00:02:01,792
all of these different factors are going to work together,
62

62

00:02:01,792  -->  00:02:04,710
just like gears, and each one turns another,
63

63

00:02:04,710  -->  00:02:06,750
and each one is going to feed another one,
64

64

00:02:06,750  -->  00:02:08,650
and the more reliability and availability
65

65

00:02:08,650  -->  00:02:09,530
you get in your networks
66

66

00:02:09,530  -->  00:02:11,910
by adding all these components together.
67

67

00:02:11,910  -->  00:02:12,880
In addition to all this,
68

68

00:02:12,880  -->  00:02:13,968
there are other design considerations
69

69

00:02:13,968  -->  00:02:16,090
that we have to think about as well,
70

70

00:02:16,090  -->  00:02:18,110
like what redundancy features should we use
71

71

00:02:18,110  -->  00:02:21,000
in terms of powering the infrastructure devices?
72

72

00:02:21,000  -->  00:02:22,550
Are we going to have internal power supplies
73

73

00:02:22,550  -->  00:02:24,581
and have two of those, and have them redundant?
74

74

00:02:24,581  -->  00:02:27,370
Or, are we going to have battery backups, or UPSs,
75

75

00:02:27,370  -->  00:02:28,850
are we going to have generators?
76

76

00:02:28,850  -->  00:02:30,970
All of these things are things you have to think about,
77

77

00:02:30,970  -->  00:02:33,080
and I don't have necessarily the right answers for you,
78

78

00:02:33,080  -->  00:02:36,030
because it all comes down to a case-by-case basis.
79

79

00:02:36,030  -->  00:02:37,750
Every network is going to be different,
80

80

00:02:37,750  -->  00:02:39,270
and every one has its own needs
81

81

00:02:39,270  -->  00:02:41,730
and its own business case associated with it.
82

82

00:02:41,730  -->  00:02:43,255
The network that I had at former employers
83

83

00:02:43,255  -->  00:02:45,650
were serving hundreds of thousands of clients,
84

84

00:02:45,650  -->  00:02:47,710
and those were vastly different than the ones
85

85

00:02:47,710  -->  00:02:49,850
that are servicing my training company right now,
86

86

00:02:49,850  -->  00:02:51,650
with just a handful of employees.
87

87

00:02:51,650  -->  00:02:53,500
Because when you're dealing with your network design
88

88

00:02:53,500  -->  00:02:54,600
and your redundancies,
89

89

00:02:54,600  -->  00:02:57,320
you have to think about the business case first.
90

90

00:02:57,320  -->  00:02:58,540
Each one is going to be different
91

91

00:02:58,540  -->  00:03:01,160
based on your needs and your considerations.
92

92

00:03:01,160  -->  00:03:02,850
What redundancy features should be used
93

93

00:03:02,850  -->  00:03:05,310
to maintain the environmental conditions of your space?
94

94

00:03:05,310  -->  00:03:07,580
If you have good power and space and cooling,
95

95

00:03:07,580  -->  00:03:08,413
you need to make sure
96

96

00:03:08,413  -->  00:03:09,820
that you're thinking about air conditioning,
97

97

00:03:09,820  -->  00:03:11,650
and do you have one unit or two?
98

98

00:03:11,650  -->  00:03:13,210
Do you have generators onsite?
99

99

00:03:13,210  -->  00:03:16,080
Do you have additional thermal heating or thermal cooling?
100

100

00:03:16,080  -->  00:03:18,200
All of these things are things you have to think about.
101

101

00:03:18,200  -->  00:03:20,050
What do you do when power goes down?
102

102

00:03:20,050  -->  00:03:21,020
What are some of those things
103

103

00:03:21,020  -->  00:03:21,860
that you're going to have to deal with
104

104

00:03:21,860  -->  00:03:23,290
if you're running a server farm
105

105

00:03:23,290  -->  00:03:25,760
that has to have units running all the time,
106

106

00:03:25,760  -->  00:03:27,160
because it can't afford to go down
107

107

00:03:27,160  -->  00:03:28,220
because it's going to affect
108

108

00:03:28,220  -->  00:03:29,930
thousands and thousands of people,
109

109

00:03:29,930  -->  00:03:32,540
instead of just your one office with 20 people?
110

110

00:03:32,540  -->  00:03:34,140
All of these are things you have to consider
111

111

00:03:34,140  -->  00:03:35,370
as you think about it.
112

112

00:03:35,370  -->  00:03:37,310
In my office, we made the decision
113

113

00:03:37,310  -->  00:03:39,380
that one air conditioning unit was enough,
114

114

00:03:39,380  -->  00:03:41,930
because if it goes down, we might just not work today
115

115

00:03:41,930  -->  00:03:44,510
and we'll come to work tomorrow, we can get over that.
116

116

00:03:44,510  -->  00:03:45,940
But in a server farm,
117

117

00:03:45,940  -->  00:03:48,310
we need to make sure we have multiple air conditioners,
118

118

00:03:48,310  -->  00:03:49,143
because if that goes down
119

119

00:03:49,143  -->  00:03:51,240
it can actually burn up all the components, right?
120

120

00:03:51,240  -->  00:03:53,220
So we have to have additional power and space and cooling
121

121

00:03:53,220  -->  00:03:54,840
that are fully redundant,
122

122

00:03:54,840  -->  00:03:56,226
because of that server infrastructure
123

123

00:03:56,226  -->  00:03:58,040
that we're supporting there.
124

124

00:03:58,040  -->  00:04:00,930
These are the things you have to balance in your practices.
125

125

00:04:00,930  -->  00:04:02,807
And so when you start looking at the best practices,
126

126

00:04:02,807  -->  00:04:05,120
I want you to examine your technical goals
127

127

00:04:05,120  -->  00:04:06,780
and your operational goals.
128

128

00:04:06,780  -->  00:04:08,150
Now what I mean by that is,
129

129

00:04:08,150  -->  00:04:10,230
what is the function of this network?
130

130

00:04:10,230  -->  00:04:12,200
What are you actually trying to accomplish?
131

131

00:04:12,200  -->  00:04:16,350
Are you trying to get to 90% uptime, or 95%, or 99%,
132

132

00:04:16,350  -->  00:04:17,677
or are you going for that gold standard
133

133

00:04:17,677  -->  00:04:20,100
of five nines of availability?
134

134

00:04:20,100  -->  00:04:22,200
Every company has a different technical goal,
135

135

00:04:22,200  -->  00:04:23,960
and that technical goal is going to determine
136

136

00:04:23,960  -->  00:04:25,470
the design of your network.
137

137

00:04:25,470  -->  00:04:26,519
And you need to identify that
138

138

00:04:26,519  -->  00:04:28,390
inside of your budgeting as well,
139

139

00:04:28,390  -->  00:04:30,471
because funding these high-availability features
140

140

00:04:30,471  -->  00:04:32,440
is really expensive.
141

141

00:04:32,440  -->  00:04:34,630
As I said, if I want to put a second router in there,
142

142

00:04:34,630  -->  00:04:37,820
that might cost me another 3,000 or $5,000.
143

143

00:04:37,820  -->  00:04:39,160
In my own personal network,
144

144

00:04:39,160  -->  00:04:42,050
we have a file server, and it's a small NAS device.
145

145

00:04:42,050  -->  00:04:45,150
We're not comfortable just having all of our devices there,
146

146

00:04:45,150  -->  00:04:46,480
so we decided we weren't comfortable
147

147

00:04:46,480  -->  00:04:48,980
having all of our file storage on a single hard drive,
148

148

00:04:48,980  -->  00:04:51,070
and we built this NAS array instead,
149

149

00:04:51,070  -->  00:04:52,760
so if one of those drives goes out,
150

150

00:04:52,760  -->  00:04:55,220
we have three others that are carrying the load.
151

151

00:04:55,220  -->  00:04:56,590
This is the idea here.
152

152

00:04:56,590  -->  00:04:59,300
Now, eventually we decided we didn't need that NAS anymore,
153

153

00:04:59,300  -->  00:05:02,550
and so we replaced that NAS enclosure with a full RAID 5.
154

154

00:05:02,550  -->  00:05:04,320
Later on we took that full RAID 5
155

155

00:05:04,320  -->  00:05:06,180
and we switched it over to a cloud server
156

156

00:05:06,180  -->  00:05:07,310
that has redundant backups
157

157

00:05:07,310  -->  00:05:09,040
in two different cloud environments.
158

158

00:05:09,040  -->  00:05:10,740
And so all of these things work together
159

159

00:05:10,740  -->  00:05:12,100
based on our decisions,
160

160

00:05:12,100  -->  00:05:13,560
but as we moved up that scale
161

161

00:05:13,560  -->  00:05:15,180
and got more and more redundancy,
162

162

00:05:15,180  -->  00:05:17,150
we have more and more costs associated.
163

163

00:05:17,150  -->  00:05:19,440
It was a lot cheaper just to have an 8-terabyte hard drive
164

164

00:05:19,440  -->  00:05:20,780
with all of our files on it,
165

165

00:05:20,780  -->  00:05:21,960
then we went to a NAS array
166

166

00:05:21,960  -->  00:05:23,670
and that cost two or three times that money,
167

167

00:05:23,670  -->  00:05:24,586
then we went to a full RAID 5
168

168

00:05:24,586  -->  00:05:26,640
and that cost a couple more times that,
169

169

00:05:26,640  -->  00:05:29,010
then we went to the cloud and we have to pay more for that.
170

170

00:05:29,010  -->  00:05:30,740
Remember, all your decisions here
171

171

00:05:30,740  -->  00:05:32,350
are going to cost you more money,
172

172

00:05:32,350  -->  00:05:35,440
but if it's worth it to you, that would be important, right,
173

173

00:05:35,440  -->  00:05:37,071
and so these are the things you have to balance
174

174

00:05:37,071  -->  00:05:39,590
as you're designing these fully redundant networks,
175

175

00:05:39,590  -->  00:05:41,800
based on those technical goals.
176

176

00:05:41,800  -->  00:05:42,970
You also need to categorize
177

177

00:05:42,970  -->  00:05:45,630
all of your business applications into profiles,
178

178

00:05:45,630  -->  00:05:47,180
to help with this redundancy mission
179

179

00:05:47,180  -->  00:05:49,070
that you're trying to go and accomplish here.
180

180

00:05:49,070  -->  00:05:50,280
This will really help you as you start
181

181

00:05:50,280  -->  00:05:52,630
going into the quality of service as well.
182

182

00:05:52,630  -->  00:05:53,900
Now if I said, for instance,
183

183

00:05:53,900  -->  00:05:55,940
that web is considered category one
184

184

00:05:55,940  -->  00:05:57,550
and email is category two
185

185

00:05:57,550  -->  00:05:59,580
and streaming video's going to be category three,
186

186

00:05:59,580  -->  00:06:00,732
then we can apply profiles
187

187

00:06:00,732  -->  00:06:02,291
and give certain levels of service
188

188

00:06:02,291  -->  00:06:04,290
to each of those categories.
189

189

00:06:04,290  -->  00:06:06,450
Now we'll talk specifically of how that works
190

190

00:06:06,450  -->  00:06:09,200
when we talk about quality of service in a future lesson.
191

191

00:06:09,200  -->  00:06:10,240
Another thing we want to do
192

192

00:06:10,240  -->  00:06:11,820
is establish performance standards
193

193

00:06:11,820  -->  00:06:14,010
for our high-availability networks.
194

194

00:06:14,010  -->  00:06:16,370
What are the standards that we're going to have to have?
195

195

00:06:16,370  -->  00:06:17,670
These standards are going to drive
196

196

00:06:17,670  -->  00:06:19,550
how success is measured for us,
197

197

00:06:19,550  -->  00:06:21,960
and in the case of my file server, for instance,
198

198

00:06:21,960  -->  00:06:24,540
we measure success as it being up and available
199

199

00:06:24,540  -->  00:06:26,660
when my video editors need to access it,
200

200

00:06:26,660  -->  00:06:28,020
and that they don't lose data,
201

201

00:06:28,020  -->  00:06:29,520
because if we lost all of our files,
202

202

00:06:29,520  -->  00:06:31,110
that'd be bad for us, right?
203

203

00:06:31,110  -->  00:06:32,570
Those are two metrics that we have,
204

204

00:06:32,570  -->  00:06:35,035
and we have numbers associated with each of those things.
205

205

00:06:35,035  -->  00:06:38,020
In other organizations, we measure it based on the uptime
206

206

00:06:38,020  -->  00:06:40,030
of the entire end-to-end service,
207

207

00:06:40,030  -->  00:06:42,670
so if a client can't get out to the internet for an ISP,
208

208

00:06:42,670  -->  00:06:45,530
that would be a bad thing, that's one of their measurements.
209

209

00:06:45,530  -->  00:06:47,890
Now the other one might be, what is their uptime?
210

210

00:06:47,890  -->  00:06:49,420
All of these performance standards are developed
211

211

00:06:49,420  -->  00:06:51,760
through metrics and key performance indicators.
212

212

00:06:51,760  -->  00:06:52,950
If you're using something like ITIL
213

213

00:06:52,950  -->  00:06:54,850
as your IT service management standards,
214

214

00:06:54,850  -->  00:06:56,520
this is what you're going to be doing as you're trying
215

215

00:06:56,520  -->  00:06:59,310
to run those inside your organization as well.
216

216

00:06:59,310  -->  00:07:02,490
Finally, here we wanted to find how we manage and measure
217

217

00:07:02,490  -->  00:07:04,940
the high-availability solutions for ourselves.
218

218

00:07:04,940  -->  00:07:07,710
Metrics are going to be really useful to quantify success,
219

219

00:07:07,710  -->  00:07:09,860
if you develop those metrics correctly.
220

220

00:07:09,860  -->  00:07:12,920
Decision-makers and leaders love seeing metrics.
221

221

00:07:12,920  -->  00:07:15,180
They love seeing charts and seeing the performance,
222

222

00:07:15,180  -->  00:07:16,640
and how it's going up over time,
223

223

00:07:16,640  -->  00:07:18,200
and how our availability is going up,
224

224

00:07:18,200  -->  00:07:19,970
and how our costs are going down.
225

225

00:07:19,970  -->  00:07:21,330
Those are all good things,
226

226

00:07:21,330  -->  00:07:22,780
but if you don't know what you're measuring
227

227

00:07:22,780  -->  00:07:24,110
or why you're measuring it,
228

228

00:07:24,110  -->  00:07:26,190
it really goes back to your performance standards.
229

229

00:07:26,190  -->  00:07:27,470
Then, these are the kind of things
230

230

00:07:27,470  -->  00:07:29,220
that are wasting your time with metrics.
231

231

00:07:29,220  -->  00:07:31,130
A lot of people measure a lot of things,
232

232

00:07:31,130  -->  00:07:32,140
and they don't really tell you
233

233

00:07:32,140  -->  00:07:34,240
if you're getting the outcome you're wanting.
234

234

00:07:34,240  -->  00:07:35,960
I want to make sure that you think about
235

235

00:07:35,960  -->  00:07:38,570
how you decide on what metrics you're going to use.
236

236

00:07:38,570  -->  00:07:40,478
Now, we've covered a lot of different design criteria
237

237

00:07:40,478  -->  00:07:43,448
in this lesson, but the real big takeaway here
238

238

00:07:43,448  -->  00:07:45,530
that I want you to think about is this.
239

239

00:07:45,530  -->  00:07:47,230
If you have an existing network,
240

240

00:07:47,230  -->  00:07:48,820
you can add availability to it,
241

241

00:07:48,820  -->  00:07:50,600
and you can add redundancy to it.
242

242

00:07:50,600  -->  00:07:52,320
You can retrofit stuff in,
243

243

00:07:52,320  -->  00:07:53,834
but it's going to cost you a lot more time
244

244

00:07:53,834  -->  00:07:55,720
and a lot more money.
245

245

00:07:55,720  -->  00:07:57,600
It is much, much cheaper
246

246

00:07:57,600  -->  00:07:59,377
to design this stuff early in the process
247

247

00:07:59,377  -->  00:08:01,740
when you start building a network from scratch.
248

248

00:08:01,740  -->  00:08:04,490
So, if you're designing a network and you're asked early on
249

249

00:08:04,490  -->  00:08:05,970
what kind of things you need,
250

250

00:08:05,970  -->  00:08:08,480
I want you to think about all these things of redundancy
251

251

00:08:08,480  -->  00:08:10,020
in your initial design.
252

252

00:08:10,020  -->  00:08:13,210
Adding them in early is going to save you a lot of money.
253

253

00:08:13,210  -->  00:08:15,470
Every project has three main factors,
254

254

00:08:15,470  -->  00:08:17,680
time, cost, and quality,
255

255

00:08:17,680  -->  00:08:20,029
and usually, one of these things is going to suffer
256

256

00:08:20,029  -->  00:08:22,300
at the expense of the other two.
257

257

00:08:22,300  -->  00:08:24,570
For example, if I asked you to build me a network
258

258

00:08:24,570  -->  00:08:25,940
and I want it to be fully redundant
259

259

00:08:25,940  -->  00:08:28,600
and available by tomorrow, could you do it?
260

260

00:08:28,600  -->  00:08:32,000
Well, maybe, but it's probably going to cost me a lot of money,
261

261

00:08:32,000  -->  00:08:33,790
and because they give you very little time,
262

262

00:08:33,790  -->  00:08:35,340
it's going to cost me even more,
263

263

00:08:35,340  -->  00:08:37,420
or your quality is going to suffer.
264

264

00:08:37,420  -->  00:08:39,810
So, you could do it good, you could do it quick,
265

265

00:08:39,810  -->  00:08:42,610
or you could do it cheap, but you can't do all three.
266

266

00:08:42,610  -->  00:08:45,260
It's always going to be a trade-off between these three things,
267

267

00:08:45,260  -->  00:08:46,270
and I want you to remember
268

268

00:08:46,270  -->  00:08:48,610
as you're out there and you're designing networks,
269

269

00:08:48,610  -->  00:08:50,850
you need to make sure you're thinking about your redundancy
270

270

00:08:50,850  -->  00:08:53,190
and your availability and your reliability,
271

271

00:08:53,190  -->  00:08:55,393
because often that quality is going to suffer
272

272

00:08:55,393  -->  00:08:57,630
in favor of getting things out quicker
273

273

00:08:57,630  -->  00:08:59,130
or getting things out cheaper.
