1
1

00:00:00,630  -->  00:00:02,670
<v Instructor>Security Data Collection.</v>
2

2

00:00:02,670  -->  00:00:05,160
In this lesson, we're going to talk about all that data
3

3

00:00:05,160  -->  00:00:07,770
that you're collecting inside your SIEM.
4

4

00:00:07,770  -->  00:00:09,720
Now, a lot of this can become intelligence,
5

5

00:00:09,720  -->  00:00:13,110
but intelligence loses its value over time.
6

6

00:00:13,110  -->  00:00:15,030
So when you're dealing with this, you need to make sure
7

7

00:00:15,030  -->  00:00:17,310
that you're capturing and analyzing the information
8

8

00:00:17,310  -->  00:00:20,610
in real time or as close to real time as possible.
9

9

00:00:20,610  -->  00:00:21,870
The sooner you can find out
10

10

00:00:21,870  -->  00:00:23,820
about a bad guy intruding into your network,
11

11

00:00:23,820  -->  00:00:25,440
the quicker you can get them out, right?
12

12

00:00:25,440  -->  00:00:28,710
And so intelligence needs to be current and relevant.
13

13

00:00:28,710  -->  00:00:31,590
Now, as we talked about the intelligence process earlier,
14

14

00:00:31,590  -->  00:00:34,470
we talked about the fact that we have five different stages.
15

15

00:00:34,470  -->  00:00:35,970
We start out with Requirements,
16

16

00:00:35,970  -->  00:00:37,800
we move into Collection &amp; Processing,
17

17

00:00:37,800  -->  00:00:40,980
then Analysis, then Dissemination, and then Feedback.
18

18

00:00:40,980  -->  00:00:42,120
Here on the screen,
19

19

00:00:42,120  -->  00:00:43,560
you can see the fact that I've highlighted
20

20

00:00:43,560  -->  00:00:46,740
Collection &amp; Processing, Analysis, and Dissemination.
21

21

00:00:46,740  -->  00:00:47,760
The reason for that
22

22

00:00:47,760  -->  00:00:50,130
is this is what we're focusing on right now.
23

23

00:00:50,130  -->  00:00:51,810
When we're talking about a SIEM,
24

24

00:00:51,810  -->  00:00:53,520
this is where a SIEM operates.
25

25

00:00:53,520  -->  00:00:55,080
It helps you collect the information.
26

26

00:00:55,080  -->  00:00:57,750
It helps you process that information and normalize it.
27

27

00:00:57,750  -->  00:00:59,430
It helps you analyze that information.
28

28

00:00:59,430  -->  00:01:01,380
And then you can even run reports
29

29

00:01:01,380  -->  00:01:04,560
or send information out to others as part of dissemination.
30

30

00:01:04,560  -->  00:01:05,850
And so a SIEM really does fit
31

31

00:01:05,850  -->  00:01:08,910
into these three phases of the intelligence lifecycle.
32

32

00:01:08,910  -->  00:01:11,100
Now, one of the things your SIEMs can do for you is,
33

33

00:01:11,100  -->  00:01:12,210
you can actually configure them
34

34

00:01:12,210  -->  00:01:15,270
to automate much of the security intelligence lifecycle,
35

35

00:01:15,270  -->  00:01:16,530
especially when it comes
36

36

00:01:16,530  -->  00:01:19,110
to gathering and collecting that data.
37

37

00:01:19,110  -->  00:01:21,000
I don't want to have to go out to all these different systems
38

38

00:01:21,000  -->  00:01:22,860
across my network and grab their logs
39

39

00:01:22,860  -->  00:01:25,110
and start analyzing them, but using a SIEM,
40

40

00:01:25,110  -->  00:01:27,870
I can make all those systems feed me that data,
41

41

00:01:27,870  -->  00:01:29,730
and it can go into that central repository
42

42

00:01:29,730  -->  00:01:31,650
that we can later analyze.
43

43

00:01:31,650  -->  00:01:33,540
Now, one of the big things you have to consider
44

44

00:01:33,540  -->  00:01:35,160
when you're configuring your SIEMs is,
45

45

00:01:35,160  -->  00:01:36,690
what do you want to collect?
46

46

00:01:36,690  -->  00:01:38,190
Because some people have a tendency
47

47

00:01:38,190  -->  00:01:39,780
to just try to collect everything,
48

48

00:01:39,780  -->  00:01:41,940
they dump all the data into the system.
49

49

00:01:41,940  -->  00:01:43,140
But the problem with that
50

50

00:01:43,140  -->  00:01:45,570
is it can end up overloading your system.
51

51

00:01:45,570  -->  00:01:47,220
All this data has to be stored,
52

52

00:01:47,220  -->  00:01:49,440
it has to be processed, and it has to be normalized.
53

53

00:01:49,440  -->  00:01:51,870
And if I'm sending it in with millions of endpoints,
54

54

00:01:51,870  -->  00:01:54,120
that can really quickly overwhelm my systems.
55

55

00:01:54,120  -->  00:01:56,460
So instead, you should spend some time upfront
56

56

00:01:56,460  -->  00:01:58,950
doing your planning based on the requirements
57

57

00:01:58,950  -->  00:02:01,170
and determining exactly what you need to collect.
58

58

00:02:01,170  -->  00:02:04,080
Remember, while your SIEM could collect all the logs
59

59

00:02:04,080  -->  00:02:06,900
across all of your systems, this isn't a good idea.
60

60

00:02:06,900  -->  00:02:09,270
Instead, you need to configure your SIEM to focus
61

61

00:02:09,270  -->  00:02:12,510
on the events related to the things that you need to know.
62

62

00:02:12,510  -->  00:02:13,830
Not everything is important,
63

63

00:02:13,830  -->  00:02:16,980
so you need to identify what is and collect on that.
64

64

00:02:16,980  -->  00:02:18,930
Now, one of the biggest features of a SIEM
65

65

00:02:18,930  -->  00:02:20,670
is the ability to process data
66

66

00:02:20,670  -->  00:02:23,250
and then look for different trends and alert on those.
67

67

00:02:23,250  -->  00:02:24,990
Now, just like all alerting systems,
68

68

00:02:24,990  -->  00:02:26,430
it does suffer from the problems
69

69

00:02:26,430  -->  00:02:28,710
of false positives and false negatives.
70

70

00:02:28,710  -->  00:02:30,870
When we talk about the problem with false negatives,
71

71

00:02:30,870  -->  00:02:32,730
this is when security administrators are exposed
72

72

00:02:32,730  -->  00:02:34,560
to a threat without being aware of them
73

73

00:02:34,560  -->  00:02:37,710
because your system falsely categorized it as negative
74

74

00:02:37,710  -->  00:02:40,200
instead of alerting that there was something bad there.
75

75

00:02:40,200  -->  00:02:41,910
Now, on the other hand, we also have issues
76

76

00:02:41,910  -->  00:02:43,380
when we have false positives.
77

77

00:02:43,380  -->  00:02:45,300
If our system starts having a lot of positives
78

78

00:02:45,300  -->  00:02:46,440
but they aren't real,
79

79

00:02:46,440  -->  00:02:48,930
we're going to overwhelm our analysis and response resources
80

80

00:02:48,930  -->  00:02:50,880
because some person has to look at that
81

81

00:02:50,880  -->  00:02:52,200
and analyze it and determine,
82

82

00:02:52,200  -->  00:02:54,030
was there really an event that happened?
83

83

00:02:54,030  -->  00:02:55,800
So we want to make sure we are tuning our systems
84

84

00:02:55,800  -->  00:02:56,970
to make sure that we don't have
85

85

00:02:56,970  -->  00:02:59,520
a lot of false negatives or false positives.
86

86

00:02:59,520  -->  00:03:02,610
To help us do that, we develop what's called a use case.
87

87

00:03:02,610  -->  00:03:03,960
By developing use cases,
88

88

00:03:03,960  -->  00:03:07,050
we can mitigate the risk of these false indicators.
89

89

00:03:07,050  -->  00:03:08,520
Now, when I talk about a use case,
90

90

00:03:08,520  -->  00:03:11,250
this is a specific condition that should be reported,
91

91

00:03:11,250  -->  00:03:13,050
such as a suspicious log on,
92

92

00:03:13,050  -->  00:03:15,270
or a process executing from a temporary directory,
93

93

00:03:15,270  -->  00:03:16,740
or something like that.
94

94

00:03:16,740  -->  00:03:18,990
Essentially, we want to think about what is this bad thing
95

95

00:03:18,990  -->  00:03:20,280
that we want to collect on?
96

96

00:03:20,280  -->  00:03:22,470
And then we develop this use case around it.
97

97

00:03:22,470  -->  00:03:23,580
Based on that use case,
98

98

00:03:23,580  -->  00:03:26,970
we can then configure our SIEM to collect the relevant data.
99

99

00:03:26,970  -->  00:03:28,470
Now, what we want to do here is essentially
100

100

00:03:28,470  -->  00:03:30,960
develop a template for each of these use cases.
101

101

00:03:30,960  -->  00:03:32,160
And as we do that,
102

102

00:03:32,160  -->  00:03:34,470
they're going to contain a couple of different things.
103

103

00:03:34,470  -->  00:03:36,360
We're going to contain things like the data sources
104

104

00:03:36,360  -->  00:03:38,580
with the indicators that we want to collect on.
105

105

00:03:38,580  -->  00:03:40,080
It's going to contain the query strings
106

106

00:03:40,080  -->  00:03:42,120
that we're going to use to correlate those different indicators
107

107

00:03:42,120  -->  00:03:43,620
across different systems.
108

108

00:03:43,620  -->  00:03:45,060
We want to make sure we have the actions
109

109

00:03:45,060  -->  00:03:47,040
that are going to occur when the event is triggered.
110

110

00:03:47,040  -->  00:03:49,080
Essentially, what are we going to do to respond
111

111

00:03:49,080  -->  00:03:50,790
when we see this bad thing happen?
112

112

00:03:50,790  -->  00:03:52,260
And then by having those three things
113

113

00:03:52,260  -->  00:03:53,940
as part of our use case template,
114

114

00:03:53,940  -->  00:03:55,800
that tells us that this set
115

115

00:03:55,800  -->  00:03:57,930
is what is known as this bad thing.
116

116

00:03:57,930  -->  00:03:59,850
And so eventually, we would write a rule
117

117

00:03:59,850  -->  00:04:01,530
or a query to be able to identify
118

118

00:04:01,530  -->  00:04:03,960
all of those things across our systems.
119

119

00:04:03,960  -->  00:04:06,390
In addition to providing those three things in the use case,
120

120

00:04:06,390  -->  00:04:08,520
we also need to make sure that each use case
121

121

00:04:08,520  -->  00:04:11,280
captures the 5Ws when we're dealing with an event.
122

122

00:04:11,280  -->  00:04:14,250
This would be things like when, when did this event start?
123

123

00:04:14,250  -->  00:04:17,040
And when did this event end, if it's already ended?
124

124

00:04:17,040  -->  00:04:18,810
We also want to figure out who,
125

125

00:04:18,810  -->  00:04:22,500
who was involved in this event, which user or which system?
126

126

00:04:22,500  -->  00:04:24,960
And then we want to figure out what, what happened
127

127

00:04:24,960  -->  00:04:27,750
and what is the specific details of this event?
128

128

00:04:27,750  -->  00:04:29,610
Essentially, did somebody try to run a program
129

129

00:04:29,610  -->  00:04:30,810
and there was malware in it?
130

130

00:04:30,810  -->  00:04:32,880
Did somebody try to attack our network from the outside?
131

131

00:04:32,880  -->  00:04:35,790
What happened? We need those specific details.
132

132

00:04:35,790  -->  00:04:39,030
Then we go and figure out where, where did the event happen?
133

133

00:04:39,030  -->  00:04:43,020
Was it on a host, a server, a file system, the network?
134

134

00:04:43,020  -->  00:04:44,730
Where is this issue?
135

135

00:04:44,730  -->  00:04:46,650
And then we want to also figure out where,
136

136

00:04:46,650  -->  00:04:48,660
where did the event originate from?
137

137

00:04:48,660  -->  00:04:51,360
Did it come from the inside 'cause it was an insider threat?
138

138

00:04:51,360  -->  00:04:52,350
Did it come from the outside
139

139

00:04:52,350  -->  00:04:54,030
because there was an external hacker?
140

140

00:04:54,030  -->  00:04:55,410
This is an important piece of information
141

141

00:04:55,410  -->  00:04:56,460
for us to know too.
142

142

00:04:56,460  -->  00:04:58,110
So by knowing those 5Ws,
143

143

00:04:58,110  -->  00:04:59,580
we're going to be able to better understand
144

144

00:04:59,580  -->  00:05:01,980
what happened and then how we can respond to it.
