1
00:00:00,210 --> 00:00:02,190
Investigative data.

2
00:00:02,190 --> 00:00:03,930
In this lesson, we're going to talk about some

3
00:00:03,930 --> 00:00:06,360
of the different pieces of information you may use

4
00:00:06,360 --> 00:00:08,850
when you're conducting an incident response.

5
00:00:08,850 --> 00:00:11,340
Now, for the exam, you don't need to be an expert

6
00:00:11,340 --> 00:00:13,320
in all of these different tools or areas

7
00:00:13,320 --> 00:00:15,300
that we're going to talk about, but instead,

8
00:00:15,300 --> 00:00:16,800
you need to be able to understand

9
00:00:16,800 --> 00:00:18,870
what sources are available to you,

10
00:00:18,870 --> 00:00:21,930
and pick the right one based on a given scenario.

11
00:00:21,930 --> 00:00:23,700
So as we look at this, you have to remember

12
00:00:23,700 --> 00:00:24,870
that when we look at a network,

13
00:00:24,870 --> 00:00:27,780
there are tons and tons of pieces of data,

14
00:00:27,780 --> 00:00:29,730
and as a security analyst, it's our job

15
00:00:29,730 --> 00:00:31,380
to take that information in

16
00:00:31,380 --> 00:00:33,660
and create a consolidated picture.

17
00:00:33,660 --> 00:00:35,460
Based on that picture, we're going to be able

18
00:00:35,460 --> 00:00:38,220
to understand better what is happening inside our network

19
00:00:38,220 --> 00:00:41,100
and in response to the event we're working on.

20
00:00:41,100 --> 00:00:43,380
Now, the first thing we're going to talk about is a SIEM.

21
00:00:43,380 --> 00:00:45,090
Now, we've talked about SIEMs before,

22
00:00:45,090 --> 00:00:47,040
but a SIEM is a Security Information

23
00:00:47,040 --> 00:00:48,690
and Event Monitoring System.

24
00:00:48,690 --> 00:00:51,180
Now, this is important because it's going to be a combination

25
00:00:51,180 --> 00:00:54,930
of a lot of different data sources into this one SIEM tool,

26
00:00:54,930 --> 00:00:56,940
and this provides us with real-time analysis

27
00:00:56,940 --> 00:00:58,650
of security alerts that are generated

28
00:00:58,650 --> 00:01:01,290
by applications and network hardware.

29
00:01:01,290 --> 00:01:02,670
As we go through, we're going to be able

30
00:01:02,670 --> 00:01:04,530
to see a SIEM dashboard.

31
00:01:04,530 --> 00:01:06,360
Now, from that dashboard, we could see the number

32
00:01:06,360 --> 00:01:07,560
of hosts we have, the number of authentications

33
00:01:07,560 --> 00:01:10,620
that have happened, the number of unique IPs we're seeing.

34
00:01:10,620 --> 00:01:12,360
All the different hosts in our network,

35
00:01:12,360 --> 00:01:13,320
and we can cut through

36
00:01:13,320 --> 00:01:15,390
that data using different search parameters

37
00:01:15,390 --> 00:01:17,250
to find the information we need.

38
00:01:17,250 --> 00:01:18,630
When you're doing an incident response,

39
00:01:18,630 --> 00:01:21,420
a SIEM is a really helpful thing for you.

40
00:01:21,420 --> 00:01:22,740
Now, when you think about a SIEM,

41
00:01:22,740 --> 00:01:24,960
there are lots of different pieces of information

42
00:01:24,960 --> 00:01:27,510
and a lot of ways for this information to get there.

43
00:01:27,510 --> 00:01:29,790
The first thing we have to think about is our sensor.

44
00:01:29,790 --> 00:01:32,460
This is the actual endpoint that's being monitored.

45
00:01:32,460 --> 00:01:35,700
That sensor can then feed that data up into the SIEM.

46
00:01:35,700 --> 00:01:36,600
Another thing we have to think about

47
00:01:36,600 --> 00:01:38,550
with our SIEMs is their sensitivity.

48
00:01:38,550 --> 00:01:41,250
Now, the sensitivity is focused on how much

49
00:01:41,250 --> 00:01:43,320
or how little you're going to be logging.

50
00:01:43,320 --> 00:01:45,360
Based on how you configure that sensor,

51
00:01:45,360 --> 00:01:46,800
that's going to determine how much data

52
00:01:46,800 --> 00:01:48,210
is being sent to the SIEM.

53
00:01:48,210 --> 00:01:49,920
Now, you may think it's great to send everything

54
00:01:49,920 --> 00:01:52,170
to the SIEM, and in a lot of cases it is,

55
00:01:52,170 --> 00:01:54,630
but you have to remember that a SIEM can become overloaded

56
00:01:54,630 --> 00:01:56,340
with too much information.

57
00:01:56,340 --> 00:01:58,530
All that information takes processing power,

58
00:01:58,530 --> 00:02:01,170
it takes network bandwidth, and it takes storage to hold it.

59
00:02:01,170 --> 00:02:02,310
So you have to think about these things

60
00:02:02,310 --> 00:02:03,810
as you're configuring your system

61
00:02:03,810 --> 00:02:05,760
and creating the right sensitivity levels.

62
00:02:05,760 --> 00:02:07,710
Another thing we have to think about is trends.

63
00:02:07,710 --> 00:02:09,840
By using a SIEM and its graphical ability

64
00:02:09,840 --> 00:02:11,220
to look across these logs,

65
00:02:11,220 --> 00:02:13,440
we can start seeing trends in our network.

66
00:02:13,440 --> 00:02:14,970
For instance, we may see the number

67
00:02:14,970 --> 00:02:17,160
of failed authentication attempts going up.

68
00:02:17,160 --> 00:02:18,270
That might be the indication

69
00:02:18,270 --> 00:02:20,280
that somebody's trying to brute force our network.

70
00:02:20,280 --> 00:02:22,620
These trends are very useful to us.

71
00:02:22,620 --> 00:02:24,690
Another thing we think about is the alerts.

72
00:02:24,690 --> 00:02:26,310
Inside the SIEM, we can set it up,

73
00:02:26,310 --> 00:02:27,240
so that there's certain alerts

74
00:02:27,240 --> 00:02:29,580
that happen based on certain parameters.

75
00:02:29,580 --> 00:02:31,980
For example, using the failed login attempts,

76
00:02:31,980 --> 00:02:33,540
we can use that as a good example.

77
00:02:33,540 --> 00:02:34,800
We might say that every time

78
00:02:34,800 --> 00:02:36,540
there is five failed login attempts,

79
00:02:36,540 --> 00:02:38,730
I want to have an alert sent to a system administrator

80
00:02:38,730 --> 00:02:40,140
to look into that account.

81
00:02:40,140 --> 00:02:41,910
That would be an example of an alert based

82
00:02:41,910 --> 00:02:44,070
on different inputs across the SIEM.

83
00:02:44,070 --> 00:02:45,690
And then finally, correlation.

84
00:02:45,690 --> 00:02:47,490
This is one of the big things within a SIEM

85
00:02:47,490 --> 00:02:49,230
because we're getting data from all sorts

86
00:02:49,230 --> 00:02:51,330
of different sources across all different types

87
00:02:51,330 --> 00:02:53,130
of hosts and network devices.

88
00:02:53,130 --> 00:02:54,930
All these things need to be correlated,

89
00:02:54,930 --> 00:02:57,870
so that we have a good picture of what is really happening.

90
00:02:57,870 --> 00:02:59,850
This includes making sure that that the IPs

91
00:02:59,850 --> 00:03:02,490
and host names are all using the same format.

92
00:03:02,490 --> 00:03:04,380
This also might be things like time.

93
00:03:04,380 --> 00:03:06,570
If one system is using universal time,

94
00:03:06,570 --> 00:03:08,460
another one is using Greenwich Mean time,

95
00:03:08,460 --> 00:03:09,870
another one's using European time,

96
00:03:09,870 --> 00:03:11,640
and another one's using New York time.

97
00:03:11,640 --> 00:03:13,410
That can actually be a big issue for us.

98
00:03:13,410 --> 00:03:15,840
So we want to correlate those all to a common standard,

99
00:03:15,840 --> 00:03:19,860
and normally, we do that using UTC, Universal Time.

100
00:03:19,860 --> 00:03:22,440
Now, the next thing we want to talk about is log files.

101
00:03:22,440 --> 00:03:23,850
When we talk about log files,

102
00:03:23,850 --> 00:03:26,430
this is any file that records either events that occur

103
00:03:26,430 --> 00:03:29,220
in an operating system, or other software that's running,

104
00:03:29,220 --> 00:03:31,440
or messages between different users

105
00:03:31,440 --> 00:03:33,240
of a communication software.

106
00:03:33,240 --> 00:03:35,520
Essentially, we're going to write something down.

107
00:03:35,520 --> 00:03:36,840
Now, we're going to digitally write it down,

108
00:03:36,840 --> 00:03:38,610
and that's what a log file is.

109
00:03:38,610 --> 00:03:40,950
Now, there's lots of different types of log files out there.

110
00:03:40,950 --> 00:03:42,990
For example, we have network log files

111
00:03:42,990 --> 00:03:44,670
that are going to keep track of all the things

112
00:03:44,670 --> 00:03:46,830
that are going through our routers and our switches.

113
00:03:46,830 --> 00:03:48,210
We have system log files

114
00:03:48,210 --> 00:03:49,200
that tell us what's happening

115
00:03:49,200 --> 00:03:51,180
on an individual host or server.

116
00:03:51,180 --> 00:03:53,640
We have application log files that tell us exactly

117
00:03:53,640 --> 00:03:56,190
what each application is doing on a given system.

118
00:03:56,190 --> 00:03:57,690
We have security log files.

119
00:03:57,690 --> 00:03:59,940
They're going to monitor things like failed logins

120
00:03:59,940 --> 00:04:02,880
and login successful attempts, and other things like that.

121
00:04:02,880 --> 00:04:04,017
We have web log files,

122
00:04:04,017 --> 00:04:06,480
and this might be like your proxy server logs

123
00:04:06,480 --> 00:04:08,280
where we could see what websites have been accessed

124
00:04:08,280 --> 00:04:10,890
by your users, or if you're running a web server,

125
00:04:10,890 --> 00:04:13,440
what files are being touched by an outsider

126
00:04:13,440 --> 00:04:14,910
as they're accessing that server.

127
00:04:14,910 --> 00:04:16,529
We also have DNS logs,

128
00:04:16,529 --> 00:04:17,730
and these are going to be used to tell us

129
00:04:17,730 --> 00:04:20,279
what requests been made of that DNS server,

130
00:04:20,279 --> 00:04:21,600
so we can see who's trying to get

131
00:04:21,600 --> 00:04:24,090
what IP addresses from what domain names.

132
00:04:24,090 --> 00:04:26,010
We're also going to have authentication logs.

133
00:04:26,010 --> 00:04:28,320
This is going to tell us who has successfully logged in,

134
00:04:28,320 --> 00:04:29,940
who has successfully logged out,

135
00:04:29,940 --> 00:04:32,130
or who has failed to log in or logged out.

136
00:04:32,130 --> 00:04:34,140
It'll tell us any kind of authentication across all

137
00:04:34,140 --> 00:04:36,720
of our files, our systems, and our servers.

138
00:04:36,720 --> 00:04:38,160
We also have dump files.

139
00:04:38,160 --> 00:04:40,860
Now, dump files are when things happen to crash.

140
00:04:40,860 --> 00:04:43,380
For instance, if I have a host and it crashes,

141
00:04:43,380 --> 00:04:45,510
it can actually dump the memory contents

142
00:04:45,510 --> 00:04:47,280
to disk while it's crashing.

143
00:04:47,280 --> 00:04:50,550
That can actually be uploaded as a log file into our system

144
00:04:50,550 --> 00:04:52,860
for us to use that for analysis as well.

145
00:04:52,860 --> 00:04:54,570
We also have things like VoIP.

146
00:04:54,570 --> 00:04:57,690
Now, a lot of our systems nowadays are using VoIP as well,

147
00:04:57,690 --> 00:05:00,030
and this can be captured as part of our network log files

148
00:05:00,030 --> 00:05:03,780
or specifically, as Voiceover IP devices.

149
00:05:03,780 --> 00:05:06,300
This way we can get metadata about the calls being made,

150
00:05:06,300 --> 00:05:07,980
and if we go into the call manager,

151
00:05:07,980 --> 00:05:10,260
we can actually record calls as well.

152
00:05:10,260 --> 00:05:12,690
This is one of those things that VoIP gives us that ability

153
00:05:12,690 --> 00:05:14,550
where we can see exactly who's been calling,

154
00:05:14,550 --> 00:05:17,070
how long they were calling, and even what the contents

155
00:05:17,070 --> 00:05:19,290
of that conversation were, if we have that allowed

156
00:05:19,290 --> 00:05:21,300
by our policies and procedures.

157
00:05:21,300 --> 00:05:22,650
Now, the next thing we want to talk

158
00:05:22,650 --> 00:05:26,970
about is Syslog, Rsyslog and Syslog-ng.

159
00:05:26,970 --> 00:05:28,050
Now, all three of these

160
00:05:28,050 --> 00:05:29,430
are basically three variations

161
00:05:29,430 --> 00:05:30,870
that do the same thing.

162
00:05:30,870 --> 00:05:32,340
They all are going to permit logging

163
00:05:32,340 --> 00:05:33,540
of data from different types

164
00:05:33,540 --> 00:05:36,120
of systems into a central repository.

165
00:05:36,120 --> 00:05:38,370
One of the things our SIEM relies heavily on

166
00:05:38,370 --> 00:05:41,460
is using Syslog, or Rsyslog, or Syslog-ng

167
00:05:41,460 --> 00:05:44,250
to grab that information from all the various endpoints

168
00:05:44,250 --> 00:05:46,290
and dump it into our SIEM.

169
00:05:46,290 --> 00:05:48,960
The next tool we want to talk about is Journalctl,

170
00:05:48,960 --> 00:05:51,420
and this is actually a Linux command line utility

171
00:05:51,420 --> 00:05:52,500
that's used for querying

172
00:05:52,500 --> 00:05:54,690
and displaying logs from the journald,

173
00:05:54,690 --> 00:05:56,010
which is the journald daemon

174
00:05:56,010 --> 00:05:58,020
which is basically the logging service

175
00:05:58,020 --> 00:06:00,690
for system D on a Linux machine.

176
00:06:00,690 --> 00:06:01,590
And so, if you want to be able

177
00:06:01,590 --> 00:06:03,330
to look at the logs on a Linux machine,

178
00:06:03,330 --> 00:06:06,030
you can use journalctl to do it.

179
00:06:06,030 --> 00:06:08,280
The next one we're going to talk about is NXLog.

180
00:06:08,280 --> 00:06:11,160
Now, this is a multi-platform log management tool

181
00:06:11,160 --> 00:06:13,770
that helps us to easily identify security risks,

182
00:06:13,770 --> 00:06:16,440
policy breaches, or analyze operational problems

183
00:06:16,440 --> 00:06:19,020
in server logs, operational system logs,

184
00:06:19,020 --> 00:06:20,640
and application logs.

185
00:06:20,640 --> 00:06:22,080
Now, when you think about NXLog,

186
00:06:22,080 --> 00:06:23,070
I want you to remember that it is

187
00:06:23,070 --> 00:06:25,740
a multi-platform or cross-platform tool,

188
00:06:25,740 --> 00:06:27,390
and it's also open source.

189
00:06:27,390 --> 00:06:29,610
This also means that it has a lot of similarities

190
00:06:29,610 --> 00:06:32,040
with Rsyslog or Syslog-ng.

191
00:06:32,040 --> 00:06:34,439
So what's the difference? Well, Rsyslog

192
00:06:34,439 --> 00:06:37,620
and Syslog-ng only work on Linux and Unix systems,

193
00:06:37,620 --> 00:06:40,050
but NXLog is cross-platform, so you can use it

194
00:06:40,050 --> 00:06:42,360
on Unix, Linux, and Windows too.

195
00:06:42,360 --> 00:06:44,850
The next thing we're going to talk about is NetFlow.

196
00:06:44,850 --> 00:06:46,920
Now, NetFlow is used in networking,

197
00:06:46,920 --> 00:06:49,020
and it's a network protocol system that was created

198
00:06:49,020 --> 00:06:52,110
by Cisco, and it's going to collect active IP network traffic

199
00:06:52,110 --> 00:06:54,690
as it's flowing into or out of an interface.

200
00:06:54,690 --> 00:06:56,400
So as you start thinking about things going into

201
00:06:56,400 --> 00:06:58,200
or out of your network through the firewall

202
00:06:58,200 --> 00:06:59,370
or through a router,

203
00:06:59,370 --> 00:07:01,410
NetFlow can actually capture that information.

204
00:07:01,410 --> 00:07:02,940
Now, some of the information it captures

205
00:07:02,940 --> 00:07:05,490
is things like the point of origin, the destination,

206
00:07:05,490 --> 00:07:07,740
the volume, and the pass on the network.

207
00:07:07,740 --> 00:07:09,360
This is not a packet capture.

208
00:07:09,360 --> 00:07:11,070
We're not capturing everything,

209
00:07:11,070 --> 00:07:12,180
every single one and zero

210
00:07:12,180 --> 00:07:13,710
that's going in or out of our network.

211
00:07:13,710 --> 00:07:16,200
Instead, NetFlow is more of a summarization

212
00:07:16,200 --> 00:07:18,930
of that data that's going in and out of our network.

213
00:07:18,930 --> 00:07:20,550
It can help us with things like understanding

214
00:07:20,550 --> 00:07:21,990
who's using the most bandwidth

215
00:07:21,990 --> 00:07:23,430
or where there are traffic spikes,

216
00:07:23,430 --> 00:07:25,200
but it can't tell us exactly the file

217
00:07:25,200 --> 00:07:27,180
that went into or out of our network.

218
00:07:27,180 --> 00:07:29,850
For that, we would need to have a full packet capture.

219
00:07:29,850 --> 00:07:31,620
Now, the next one we have is SFlow,

220
00:07:31,620 --> 00:07:33,750
and this stands for Sampled Flow.

221
00:07:33,750 --> 00:07:37,050
Essentially, this was an open-source version of NetFlow

222
00:07:37,050 --> 00:07:39,750
where NetFlow is made by Cisco, and it's proprietary.

223
00:07:39,750 --> 00:07:42,000
SFlow was more of the generic version.

224
00:07:42,000 --> 00:07:44,700
It's going to provide a means for exporting truncated packets,

225
00:07:44,700 --> 00:07:46,590
as well as having an interface counter

226
00:07:46,590 --> 00:07:48,570
that is going to be used for network monitoring.

227
00:07:48,570 --> 00:07:51,660
So again, we're not going to have full-packet capture here.

228
00:07:51,660 --> 00:07:53,820
We're just going to get some of the sampled flow.

229
00:07:53,820 --> 00:07:55,590
So when we talk about SFlow, a lot of times

230
00:07:55,590 --> 00:07:57,420
what it'll do is it'll do packet captures

231
00:07:57,420 --> 00:07:59,070
where it captures one out of a hundred

232
00:07:59,070 --> 00:08:01,080
or one out of a thousand packets.

233
00:08:01,080 --> 00:08:02,610
That will help us reduce the size

234
00:08:02,610 --> 00:08:03,600
while still getting an idea

235
00:08:03,600 --> 00:08:05,730
of what's going through our networks.

236
00:08:05,730 --> 00:08:07,710
The next thing we have is IPFIX

237
00:08:07,710 --> 00:08:10,980
which is the Internet Protocol Flow Information Export.

238
00:08:10,980 --> 00:08:13,290
Now, this is a universal standard for the export

239
00:08:13,290 --> 00:08:16,020
of internet protocol flow information from your routers,

240
00:08:16,020 --> 00:08:17,880
your probes, and other devices

241
00:08:17,880 --> 00:08:19,890
that's going to be used by mediation systems,

242
00:08:19,890 --> 00:08:21,390
accounting and billing systems,

243
00:08:21,390 --> 00:08:22,860
and network management systems

244
00:08:22,860 --> 00:08:25,050
to facilitate services such as measurement,

245
00:08:25,050 --> 00:08:26,910
accounting and billing by defining

246
00:08:26,910 --> 00:08:29,340
how IP flow information is to be format

247
00:08:29,340 --> 00:08:32,429
and transferred from an exporter to a collector.

248
00:08:32,429 --> 00:08:34,020
Wow, that is a mouthful,

249
00:08:34,020 --> 00:08:36,419
and you may be wondering, what did I just say?

250
00:08:36,419 --> 00:08:39,030
Well, really, what IPFIX is used for

251
00:08:39,030 --> 00:08:41,190
is on the backend of service management.

252
00:08:41,190 --> 00:08:42,090
Let's say, for instance,

253
00:08:42,090 --> 00:08:43,710
that I was running a cell phone company,

254
00:08:43,710 --> 00:08:46,200
and I was going to charge you $10 for every gigabyte

255
00:08:46,200 --> 00:08:48,060
of data that you transfer per month.

256
00:08:48,060 --> 00:08:50,460
Well, if I'm using IPFIX, I can count up

257
00:08:50,460 --> 00:08:52,080
until I get to one gigabyte,

258
00:08:52,080 --> 00:08:54,570
pass that to the billing system in the standard format,

259
00:08:54,570 --> 00:08:57,390
and then my system can charge you that $10.

260
00:08:57,390 --> 00:08:59,400
That's what IPFIX is used for.

261
00:08:59,400 --> 00:09:01,834
Now, all three of these tools, NetFlow,

262
00:09:01,834 --> 00:09:04,770
SFlow, and IPFIX can give you a good idea

263
00:09:04,770 --> 00:09:07,350
of how much bandwidth is being used in your environment.

264
00:09:07,350 --> 00:09:09,353
You can use a lot of different tools that actually make this

265
00:09:09,353 --> 00:09:11,640
in a graphical format like you see here.

266
00:09:11,640 --> 00:09:14,190
Now here you can notice that there's one big spike.

267
00:09:14,190 --> 00:09:16,230
This tells me as I'm monitoring my bandwidth that

268
00:09:16,230 --> 00:09:18,540
that point had the largest amount of usage,

269
00:09:18,540 --> 00:09:20,340
and this was traffic outbound.

270
00:09:20,340 --> 00:09:22,290
Now, it's significantly higher than the other ones.

271
00:09:22,290 --> 00:09:23,820
So as an analyst, I might go,

272
00:09:23,820 --> 00:09:26,040
hey, something doesn't seem right there.

273
00:09:26,040 --> 00:09:28,110
Why did we just double or triple the amount

274
00:09:28,110 --> 00:09:29,670
of traffic leaving our network?

275
00:09:29,670 --> 00:09:31,440
And I can go back and pull that time,

276
00:09:31,440 --> 00:09:34,110
which in this case, says 20.11.

277
00:09:34,110 --> 00:09:35,760
And as I pull that information,

278
00:09:35,760 --> 00:09:37,320
we can then look through our SIEM

279
00:09:37,320 --> 00:09:39,000
to figure out was there one host

280
00:09:39,000 --> 00:09:40,170
that was sending a lot of data?

281
00:09:40,170 --> 00:09:42,090
If so, maybe it's been infected,

282
00:09:42,090 --> 00:09:43,890
and it's doing a data exfiltration,

283
00:09:43,890 --> 00:09:45,480
or was it a lot of hosts?

284
00:09:45,480 --> 00:09:47,250
Maybe it was just the fact that there was some kind

285
00:09:47,250 --> 00:09:48,360
of a big sale on Amazon

286
00:09:48,360 --> 00:09:50,670
and everybody logged on to start buying things,

287
00:09:50,670 --> 00:09:53,010
and so data was leaving the network with their credit card

288
00:09:53,010 --> 00:09:55,440
and their queries to try to get information.

289
00:09:55,440 --> 00:09:56,730
These are the things you have to think about

290
00:09:56,730 --> 00:09:58,740
as a cybersecurity analyst.

291
00:09:58,740 --> 00:10:01,560
Now, the next thing we're going to talk about is Metadata.

292
00:10:01,560 --> 00:10:04,950
Metadata is going to be data that describes other data.

293
00:10:04,950 --> 00:10:07,140
Basically, by providing an underlying definition

294
00:10:07,140 --> 00:10:09,900
or description by summarizing basic information

295
00:10:09,900 --> 00:10:11,730
about the data that makes finding

296
00:10:11,730 --> 00:10:15,030
and working with particular instances of data much easier.

297
00:10:15,030 --> 00:10:17,010
Essentially, when you think about metadata,

298
00:10:17,010 --> 00:10:19,320
this is data about the data.

299
00:10:19,320 --> 00:10:21,000
So the easiest way to think about this

300
00:10:21,000 --> 00:10:22,950
is something like your cell phone.

301
00:10:22,950 --> 00:10:24,540
If you think about your end of the month bill

302
00:10:24,540 --> 00:10:25,830
when you get from your cell phone,

303
00:10:25,830 --> 00:10:28,590
it will show you the day and time of each of your calls,

304
00:10:28,590 --> 00:10:31,470
the number you called, and the length of that call.

305
00:10:31,470 --> 00:10:34,560
That is metadata. It's data about the call.

306
00:10:34,560 --> 00:10:36,630
It doesn't tell you what you said on that call,

307
00:10:36,630 --> 00:10:39,210
but it could remind you that on August 8th,

308
00:10:39,210 --> 00:10:41,430
I talked to my mother for five minutes

309
00:10:41,430 --> 00:10:43,500
and that would tell me that because I could see

310
00:10:43,500 --> 00:10:45,600
the time I used it, what number I called,

311
00:10:45,600 --> 00:10:47,400
and how long that call lasted,

312
00:10:47,400 --> 00:10:49,320
but I won't know exactly what was said.

313
00:10:49,320 --> 00:10:50,910
That's the idea with metadata.

314
00:10:50,910 --> 00:10:53,970
Now, does that mean metadata isn't useful? Of course, not.

315
00:10:53,970 --> 00:10:55,230
As a cybersecurity analyst,

316
00:10:55,230 --> 00:10:57,570
metadata is extremely useful to us,

317
00:10:57,570 --> 00:10:59,340
and there's lots of different places you can look

318
00:10:59,340 --> 00:11:02,250
at this metadata to get information about things.

319
00:11:02,250 --> 00:11:04,500
For instance, you might be looking at email.

320
00:11:04,500 --> 00:11:06,060
If you had somebody who sent an email

321
00:11:06,060 --> 00:11:07,680
that was part of a phishing campaign,

322
00:11:07,680 --> 00:11:09,600
you can look at the metadata about that,

323
00:11:09,600 --> 00:11:12,000
such as the time it was sent, who sent it,

324
00:11:12,000 --> 00:11:13,500
which servers it came from,

325
00:11:13,500 --> 00:11:15,180
which servers it transited through,

326
00:11:15,180 --> 00:11:16,830
and all that information to be able

327
00:11:16,830 --> 00:11:19,380
to figure out exactly what happened with that email,

328
00:11:19,380 --> 00:11:22,020
even if you didn't look at the content of the email itself.

329
00:11:22,020 --> 00:11:24,120
You might also look at mobile metadata.

330
00:11:24,120 --> 00:11:25,980
Again, using the example of a cell phone,

331
00:11:25,980 --> 00:11:27,600
you can know how much data was transferred

332
00:11:27,600 --> 00:11:29,070
or how long the calls were,

333
00:11:29,070 --> 00:11:30,750
and who those people are calling,

334
00:11:30,750 --> 00:11:32,520
and who they're talking to.

335
00:11:32,520 --> 00:11:35,070
In the case of web, we might look at things like,

336
00:11:35,070 --> 00:11:36,480
which websites are you visiting

337
00:11:36,480 --> 00:11:38,400
and how long are you staying on them?

338
00:11:38,400 --> 00:11:39,390
This is something markers

339
00:11:39,390 --> 00:11:41,100
are learning about you all the time.

340
00:11:41,100 --> 00:11:43,290
They don't know exactly what you clicked on necessarily

341
00:11:43,290 --> 00:11:45,330
or where your eyes were looking on the screen,

342
00:11:45,330 --> 00:11:47,610
but they do know how long you were on a particular page

343
00:11:47,610 --> 00:11:49,110
before you clicked away,

344
00:11:49,110 --> 00:11:51,630
and this is metadata they use to re-target you

345
00:11:51,630 --> 00:11:54,450
and work their systems to try to sell you more stuff.

346
00:11:54,450 --> 00:11:56,760
Another thing we have is file metadata.

347
00:11:56,760 --> 00:11:59,130
When I look at a particular file, like a video,

348
00:11:59,130 --> 00:12:01,830
that has a lot of metadata associated with it too.

349
00:12:01,830 --> 00:12:03,900
Who created it? When did they create it?

350
00:12:03,900 --> 00:12:05,550
When was the last time it was watched?

351
00:12:05,550 --> 00:12:07,650
How long do people watch that file for?

352
00:12:07,650 --> 00:12:10,080
All this stuff is data about the data.

353
00:12:10,080 --> 00:12:11,130
This is all metadata,

354
00:12:11,130 --> 00:12:12,870
and it can be very helpful as you're going through

355
00:12:12,870 --> 00:12:15,270
an investigation and doing an incident response.

