WEBVTT

0:00:04.340000 --> 0:00:06.620000
 Introduction to encoding.

0:00:06.620000 --> 0:00:11.000000
 So welcome everyone to the first section
 of this course where we will

0:00:11.000000 --> 0:00:15.580000
 be taking a look at encoding and the
 primary objective for this section

0:00:15.580000 --> 0:00:21.520000
 is to get an introduction as to what
 data encoding is, why it's important

0:00:21.520000 --> 0:00:29.060000
 and how it works in relation to the web or
 the functionality of web applications.

0:00:29.060000 --> 0:00:33.900000
 And in order for us to understand this
 process, we are going to be utilizing

0:00:33.900000 --> 0:00:38.580000
 a multifaceted approach that will start
 off by getting an introduction

0:00:38.580000 --> 0:00:43.840000
 as to what character sets are and
 the evolution of character sets.

0:00:43.840000 --> 0:00:48.380000
 And I'll show you how that ties into
 character set encoding because that

0:00:48.380000 --> 0:00:53.580000
 is very, very important or rather fundamental
 to the functionality of

0:00:53.580000 --> 0:00:54.560000
 web applications.

0:00:54.560000 --> 0:00:59.160000
 And of course, we'll also be tying
 in the importance of understanding

0:00:59.160000 --> 0:01:04.480000
 encoding in the context of web application
 security testing or web application

0:01:04.480000 --> 0:01:06.080000
 penetration testing.

0:01:06.080000 --> 0:01:09.400000
 So to begin with, what is data encoding?

0:01:09.400000 --> 0:01:14.800000
 Well, encoding on the web refers to
 the process of representing data and

0:01:14.800000 --> 0:01:19.620000
 information in a structured format that
 allows it to be transmitted, stored

0:01:19.620000 --> 0:01:23.820000
 and interpreted by both
 computers and humans.

0:01:23.820000 --> 0:01:29.700000
 And encoding as a result is vitally
 important or is a vitally important

0:01:29.700000 --> 0:01:35.320000
 component in the transfer and representation
 of information on the internet.

0:01:35.320000 --> 0:01:40.200000
 So encoding essentially ensures that
 data like text, images, files and

0:01:40.200000 --> 0:01:45.420000
 multimedia can be effectively communicated
 and displayed through web technologies

0:01:45.420000 --> 0:01:49.500000
 like your web browser, for example,
 and typically involves converting

0:01:49.500000 --> 0:01:55.320000
 data from its original form or format
 into a format that is suitable for

0:01:55.320000 --> 0:02:00.780000
 digital transmission and storage while
 preserving its meaning and integrity.

0:02:00.780000 --> 0:02:06.900000
 So what encoding is is essentially the
 process of converting a resource

0:02:06.900000 --> 0:02:12.480000
 on a web server in order for it to be
 suitable for transmission over the

0:02:12.480000 --> 0:02:19.400000
 internet using the protocol like TCP,
 but it's really preparing it for

0:02:19.400000 --> 0:02:25.200000
 transmission. And it needs to factor
 in that whatever is being sent, let's

0:02:25.200000 --> 0:02:30.780000
 say an HTML file is encoded and
 can be transferred successfully.

0:02:30.780000 --> 0:02:35.740000
 But in addition to that can be decoded
 successfully on the client's browser.

0:02:35.740000 --> 0:02:39.920000
 So what that means is, is that data
 encoding in the context of the web

0:02:39.920000 --> 0:02:45.260000
 is the process of encoding a resource
 on a web server like a web page,

0:02:45.260000 --> 0:02:50.440000
 transmitting it across the internet
 using HTTP or HTTPS.

0:02:50.440000 --> 0:02:56.160000
 And then on the client side, decoding
 it and rendering it on the client's

0:02:56.160000 --> 0:03:00.080000
 browser. And the key thing that you need
 to understand here is not really

0:03:00.080000 --> 0:03:04.180000
 the actual transmission of the actual
 data or resource in question.

0:03:04.180000 --> 0:03:09.360000
 It's also ensuring that the encoding
 process does not change in any way,

0:03:09.360000 --> 0:03:14.800000
 shape or form. The original, the original
 file or resource or the integrity

0:03:14.800000 --> 0:03:18.980000
 of the file. What that means is that
 in simple words is that whatever

0:03:18.980000 --> 0:03:25.100000
 is on the web server should be rendered
 on the client side without any

0:03:25.100000 --> 0:03:29.560000
 changes made to the file as a result
 of the actual encoding process.

0:03:29.560000 --> 0:03:34.140000
 Now we'll dive deeper into why there
 is encoding, you know, in terms of

0:03:34.140000 --> 0:03:42.780000
 the process itself and why multiple
 ways or techniques that encoding is

0:03:42.780000 --> 0:03:44.820000
 implemented in the process.

0:03:44.820000 --> 0:03:49.900000
 Now that begs the question, what data
 is being encoded in the first place?

0:03:49.900000 --> 0:03:54.540000
 Before we actually take a look at that
 original form or the original format,

0:03:54.540000 --> 0:03:59.020000
 let's get an understanding as to how data
 encoding plays into web application

0:03:59.020000 --> 0:04:00.940000
 penetration testing.

0:04:00.940000 --> 0:04:05.400000
 All right, so encoding based on what
 I've already told you is an essential

0:04:05.400000 --> 0:04:14.840000
 aspect of web input validation, data
 transmission and various other attack

0:04:14.840000 --> 0:04:16.980000
 vectors that we'll be exploring.

0:04:16.980000 --> 0:04:21.140000
 And it essentially again involves in
 the case of the role of encoding

0:04:21.140000 --> 0:04:25.840000
 in web app in testing, it involves
 manipulating data or converting it

0:04:25.840000 --> 0:04:30.460000
 into a different format, often to bypass security
 measures, discover vulnerabilities

0:04:30.460000 --> 0:04:32.400000
 or execute attacks.

0:04:32.400000 --> 0:04:36.680000
 So in the context of web application penetration
 testing or web application

0:04:36.680000 --> 0:04:42.120000
 security testing, we're utilizing or
 leveraging encoding to do what I've

0:04:42.120000 --> 0:04:46.140000
 just said. Number one, you know, bypassing
 security measures, discovering

0:04:46.140000 --> 0:04:47.780000
 vulnerabilities, etc.

0:04:47.780000 --> 0:04:51.840000
 And this will become very clear as
 we approach the end of this course.

0:04:51.840000 --> 0:04:55.560000
 So the point is that encoding plays
 a crucial role in discovering and

0:04:55.560000 --> 0:04:59.800000
 understanding how a web application
 handles different types of input,

0:04:59.800000 --> 0:05:05.420000
 especially when those input fields or
 forms contain special characters,

0:05:05.420000 --> 0:05:09.000000
 binary data or unexpected sequences.

0:05:09.000000 --> 0:05:13.380000
 So by testing how the web application
 responds to encoded data, security

0:05:13.380000 --> 0:05:17.860000
 testers can uncover potential vulnerabilities
 that could lead to attacks

0:05:17.860000 --> 0:05:21.780000
 such as SQL injection attacks and more
 so cross-site scripting attacks

0:05:21.780000 --> 0:05:24.700000
 and other injection based attacks.

0:05:24.700000 --> 0:05:29.000000
 So the role of encoding on the web application
 developer side is fairly

0:05:29.000000 --> 0:05:31.000000
 simple and is multifaceted.

0:05:31.000000 --> 0:05:36.140000
 So what that means is that you have
 you have standard HTML encoding and

0:05:36.140000 --> 0:05:40.240000
 I'll explain, you know, the importance
 of that and why that is implemented.

0:05:40.240000 --> 0:05:44.700000
 We have URL encoding, which you probably
 are familiar with at this point.

0:05:44.700000 --> 0:05:47.880000
 We also have base 64 encoding
 for resources.

0:05:47.880000 --> 0:05:51.900000
 So we'll be taking a look at why these
 forms of encoding are implemented

0:05:51.900000 --> 0:05:57.180000
 and then how we can leverage encoding
 as web application penetration testers

0:05:57.180000 --> 0:06:01.360000
 to identify vulnerabilities, learn
 more about how the web application

0:06:01.360000 --> 0:06:05.840000
 works and you know how we can use encoding
 to aid in our exploitation

0:06:05.840000 --> 0:06:07.840000
 of vulnerabilities.

0:06:07.840000 --> 0:06:11.740000
 So now that we have an understanding
 as to the data as to what the data

0:06:11.740000 --> 0:06:17.060000
 encoding process is at a very high level,
 let's begin at the beginning.

0:06:17.060000 --> 0:06:22.220000
 And that is where we come into or we actually
 come face to face with character

0:06:22.220000 --> 0:06:27.040000
 sets or char sets, again, depending on
 how you pronounce it, but character

0:06:27.040000 --> 0:06:32.820000
 sets are not really tied specifically
 to web applications or the actual

0:06:32.820000 --> 0:06:36.680000
 web. Character sets have been there
 for a very long time, pretty much

0:06:36.680000 --> 0:06:39.760000
 as long as computing has been alive.

0:06:39.760000 --> 0:06:42.740000
 Now, you may be wondering to yourself if
 this is your first time encountering

0:06:42.740000 --> 0:06:45.980000
 this term, what is a character set?

0:06:45.980000 --> 0:06:48.980000
 Well, a char set or a character set.

0:06:48.980000 --> 0:06:51.880000
 So char set is short for character set.

0:06:51.880000 --> 0:06:56.600000
 A character set is essentially a collection
 of characters, symbols and

0:06:56.600000 --> 0:07:02.360000
 glyphs that are associated with unique
 numeric codes or code points.

0:07:02.360000 --> 0:07:03.900000
 All right, so what does this mean?

0:07:03.900000 --> 0:07:08.800000
 Well, if you think about it, when computing
 was created or during the

0:07:08.800000 --> 0:07:11.440000
 the inception of computing.

0:07:11.440000 --> 0:07:17.180000
 So when I say that, what I mean is that
 the interaction of human beings

0:07:17.180000 --> 0:07:23.280000
 and computers, it became very apparent
 that human beings cannot communicate

0:07:23.280000 --> 0:07:29.640000
 with computers via the computer's language
 set, if you will, or language,

0:07:29.640000 --> 0:07:31.020000
 which is binary, right?

0:07:31.020000 --> 0:07:32.880000
 So zeros and ones.

0:07:32.880000 --> 0:07:36.260000
 And we know that through our keyboard,
 when you look down at your keyboard,

0:07:36.260000 --> 0:07:40.380000
 you have, you know, all the letters
 of the English alphabet or whatever

0:07:40.380000 --> 0:07:42.940000
 language you speak or
 you're accustomed to.

0:07:42.940000 --> 0:07:47.560000
 But the point is that they had to be
 a translation layer or a translation

0:07:47.560000 --> 0:07:51.140000
 of what we can communicate.

0:07:51.140000 --> 0:07:57.160000
 So essentially converting what we can
 understand and what is, you know,

0:07:57.160000 --> 0:08:02.180000
 accessible or user accessible for humans
 to the language of computers,

0:08:02.180000 --> 0:08:06.340000
 which is where we had the invention of
 a character set, which is essentially

0:08:06.340000 --> 0:08:10.940000
 a mapping in the beginning of
 the English alphabet numbers.

0:08:10.940000 --> 0:08:16.220000
 So zero to nine specific symbols like
 the question mark, brackets, curly

0:08:16.220000 --> 0:08:18.920000
 braces, etc, into binary code.

0:08:18.920000 --> 0:08:22.940000
 So a character set essentially lays out
 all letters of the English alphabet.

0:08:22.940000 --> 0:08:27.980000
 Again, for example, both uppercase, lowercase,
 and then maps it to specific

0:08:27.980000 --> 0:08:32.500000
 binary values. So that means when you
 type on your keyboard, a lowercase

0:08:32.500000 --> 0:08:38.320000
 or uppercase, the character set or
 the encoding that is configured for

0:08:38.320000 --> 0:08:41.880000
 your operating system, again, will,
 you know, translate it accordingly

0:08:41.880000 --> 0:08:45.820000
 so that the computer can understand what
 you're typing in because computers

0:08:45.820000 --> 0:08:49.320000
 do not understand alphabet,
 they only know numbers.

0:08:49.320000 --> 0:08:52.980000
 It's sort of the same issue that was
 encountered during the early days

0:08:52.980000 --> 0:08:57.540000
 of the internet with IP addresses and
 why we created domain names, right?

0:08:57.540000 --> 0:09:00.200000
 Because domains are easier
 for humans to remember.

0:09:00.200000 --> 0:09:02.120000
 Words are easier for humans to remember.

0:09:02.120000 --> 0:09:06.460000
 So it would make sense that, you know,
 the domain Google.com would be

0:09:06.460000 --> 0:09:10.740000
 much easier to access than, you know,
 having a human being remember the

0:09:10.740000 --> 0:09:13.060000
 IP address of the Google web server.

0:09:13.060000 --> 0:09:17.680000
 So the point is that character sets
 define how textual data is mapped

0:09:17.680000 --> 0:09:20.280000
 to binary values in computing systems.

0:09:20.280000 --> 0:09:25.380000
 And they play a fundamental role in
 encoding and decoding text, allowing

0:09:25.380000 --> 0:09:29.960000
 computers to store process and display
 human readable characters from

0:09:29.960000 --> 0:09:31.980000
 different writing systems.

0:09:31.980000 --> 0:09:36.540000
 Now, examples of character sets are,
 of course, ASCII, Unicode, Latin

0:09:36.540000 --> 0:09:42.120000
 one, etc. And we'll be focusing specifically
 on ASCII and Unicode, and

0:09:42.120000 --> 0:09:46.280000
 then more specifically on the
 Unicode encoding schemas.

0:09:46.280000 --> 0:09:49.640000
 So many of you should be familiar
 with ASCII, right?

0:09:49.640000 --> 0:09:51.260000
 So what is ASCII?

0:09:51.260000 --> 0:09:57.620000
 ASCII is an code for information
 interchange.

0:09:57.620000 --> 0:09:59.840000
 ASCII is a character set.

0:09:59.840000 --> 0:10:04.080000
 It really can be thought
 of as a character set.

0:10:04.080000 --> 0:10:09.440000
 So it's a widely used character encoding
 standard that utilizes 128 characters

0:10:09.440000 --> 0:10:12.460000
 and was developed in the 1960s.

0:10:12.460000 --> 0:10:16.300000
 So again, the early days of computing,
 and it was created to represent

0:10:16.300000 --> 0:10:21.780000
 text and control characters in cube in
 computers and communication equipment.

0:10:21.780000 --> 0:10:25.360000
 So in the early days, you might have
 seen that specific computing systems

0:10:25.360000 --> 0:10:31.220000
 had the, they had the text or
 the slogan ASCII compliant.

0:10:31.220000 --> 0:10:36.060000
 What that means is that it was using ASCII
 as a base for the actual translation

0:10:36.060000 --> 0:10:38.460000
 between an alphabet.

0:10:38.460000 --> 0:10:42.180000
 And it essentially ensured that, you
 know, you could pretty much interact

0:10:42.180000 --> 0:10:44.520000
 with that computer effectively.

0:10:44.520000 --> 0:10:50.180000
 So ASCII pretty much defines a set of
 codes to represent letters, numbers,

0:10:50.180000 --> 0:10:55.220000
 punctuation, and control characters used
 in the English language and basic

0:10:55.220000 --> 0:10:59.780000
 communication. So you can already tell
 what the problem with ASCII was

0:10:59.780000 --> 0:11:01.780000
 or eventually came to be.

0:11:01.780000 --> 0:11:04.800000
 And that was the fact that it
 only supported one alphabet.

0:11:04.800000 --> 0:11:09.760000
 So as comp, as computing became more popular,
 became became more widespread,

0:11:09.760000 --> 0:11:14.120000
 you know, across the entire planet,
 there was an obvious issue.

0:11:14.120000 --> 0:11:16.520000
 There were other languages or alphabets.

0:11:16.520000 --> 0:11:20.200000
 And those needed to be mapped again,
 so that computers could understand

0:11:20.200000 --> 0:11:21.540000
 what you're typing in.

0:11:21.540000 --> 0:11:26.820000
 So as I said, ASCII primarily covers English
 characters, numbers, punctuation,

0:11:26.820000 --> 0:11:32.540000
 and control characters, using 708 bits,
 typically seven to represent each

0:11:32.540000 --> 0:11:36.600000
 character. ASCII cannot be used to display
 symbols from other languages

0:11:36.600000 --> 0:11:38.040000
 like Chinese. All right.

0:11:38.040000 --> 0:11:41.800000
 So again, as the internet in the context
 of the web now, as the internet

0:11:41.800000 --> 0:11:45.580000
 became more popular, we had individuals
 from different countries trying

0:11:45.580000 --> 0:11:49.220000
 to again access the internet
 and websites.

0:11:49.220000 --> 0:11:53.740000
 And it really, the point that I want
 to clarify here is that the main

0:11:53.740000 --> 0:11:59.380000
 issue was not that, you know, individuals
 in other countries did not understand

0:11:59.380000 --> 0:12:02.960000
 English. That was not really an issue,
 because that could be translated.

0:12:02.960000 --> 0:12:08.920000
 It was more so ensuring that web pages
 that were encoded or using the

0:12:08.920000 --> 0:12:14.840000
 ASCII character set could be processed
 or could be rendered on the browsers

0:12:14.840000 --> 0:12:19.640000
 or whatever operating system was being
 used by individuals, let's say,

0:12:19.640000 --> 0:12:20.820000
 from another country.

0:12:20.820000 --> 0:12:23.760000
 Now, you may be saying, I haven't
 encountered this issue.

0:12:23.760000 --> 0:12:26.540000
 Well, I'll actually show you that
 this issue still does exist.

0:12:26.540000 --> 0:12:29.760000
 And that's why when you take a look
 at the source code of a page with

0:12:29.760000 --> 0:12:32.000000
 it, BPH, BH, DML, etc.

0:12:32.000000 --> 0:12:36.980000
 A good practice is that the encoding,
 the character set encoding will

0:12:36.980000 --> 0:12:41.100000
 always be specified in the
 head, for example, in HTML.

0:12:41.100000 --> 0:12:46.600000
 So ASCII is not used anymore in terms
 of, you know, a primary character

0:12:46.600000 --> 0:12:50.100000
 set. What is used is Unicode.

0:12:50.100000 --> 0:12:54.580000
 So in this particular slide, you can see
 a conversion of, you know, decimal

0:12:54.580000 --> 0:13:00.480000
 values to their respective symbols, or
 you know, you can look at it whatever

0:13:00.480000 --> 0:13:06.220000
 way. But what you have is in this particular
 table, so, you know, the

0:13:06.220000 --> 0:13:10.600000
 exclamation mark, a double quote,
 pound symbol, the dollar symbol.

0:13:10.600000 --> 0:13:14.520000
 And it shows you the conversion to binary,
 in this case, the ASCII character

0:13:14.520000 --> 0:13:19.900000
 set. So the binary value here or code
 point, the hex representation, the

0:13:19.900000 --> 0:13:23.300000
 octal representation,
 and the decimal value.

0:13:23.300000 --> 0:13:29.680000
 And then, of course, the HTML, the HTML
 encoded value, and we'll get to

0:13:29.680000 --> 0:13:31.920000
 HTML encoding in the next video.

0:13:31.920000 --> 0:13:37.260000
 And then, of course, you have a description
 of the actual symbol, or,

0:13:37.260000 --> 0:13:40.340000
 you know, whatever letter
 within the alphabet.

0:13:40.340000 --> 0:13:44.600000
 But again, key thing to note here with
 ASCII is that it's not really a

0:13:44.600000 --> 0:13:46.380000
 problematic character set.

0:13:46.380000 --> 0:13:51.300000
 It's just that it did not have support
 for other languages, specifically

0:13:51.300000 --> 0:13:54.580000
 non-Latin languages.

0:13:54.580000 --> 0:13:58.560000
 And in terms of its characteristics,
 as I've already stated, the character

0:13:58.560000 --> 0:14:03.640000
 set includes a total of 128 characters,
 each represented by a unique 7

0:14:03.640000 --> 0:14:05.200000
-bit binary code.

0:14:05.200000 --> 0:14:08.960000
 These characters include uppercase and
 lowercase letters, digits, punctuation

0:14:08.960000 --> 0:14:11.560000
 marks, and some control characters.

0:14:11.560000 --> 0:14:13.140000
 It has a 7-bit encoding.

0:14:13.140000 --> 0:14:16.820000
 It also supports 8-bit, but
 7-bit is the standard.

0:14:16.820000 --> 0:14:20.960000
 So in ASCII, each character is encoded
 using 7-bits, allowing for a total

0:14:20.960000 --> 0:14:25.620000
 of 2 to the power of 7,
 possible characters.

0:14:25.620000 --> 0:14:27.140000
 That's where you have the limit.

0:14:27.140000 --> 0:14:31.840000
 The most significant bit is often used
 for parity checking in older systems.

0:14:31.840000 --> 0:14:36.240000
 In terms of standardization, ASCII was
 established as a standard by the

0:14:36.240000 --> 0:14:41.800000
 American National Standards Institute,
 or ANSI, in 1963, and later became

0:14:41.800000 --> 0:14:43.160000
 an international standard.

0:14:43.160000 --> 0:14:47.240000
 So going back to the example I highlighted
 earlier, where devices were,

0:14:47.240000 --> 0:14:50.880000
 you typically see that this is, you
 know, ANSI certified or is compliant

0:14:50.880000 --> 0:14:56.740000
 with the ANSI, sorry, the ASCII standard.


0:14:56.740000 --> 0:15:01.160000
 In terms of its other characteristics,
 the basic character set, again,

0:15:01.160000 --> 0:15:06.180000
 has the uppercase letters of the
 English alphabet mapped out.

0:15:06.180000 --> 0:15:12.200000
 And those are within the range of from
 65 to 90 in terms of the actual

0:15:12.200000 --> 0:15:16.640000
 total character set, which
 again consists of 128.

0:15:16.640000 --> 0:15:21.180000
 So lowercase, what's highlighted in the
 brackets there is the actual range

0:15:21.180000 --> 0:15:24.800000
 in terms of where they fall within
 the actual character set.

0:15:24.800000 --> 0:15:30.160000
 So lowercase letters, 97 to 122, in
 terms of digits, numerical digits,

0:15:30.160000 --> 0:15:32.740000
 they fall between 48 and 57.

0:15:32.740000 --> 0:15:36.020000
 Punctuation, it has support for
 the following symbols here.

0:15:36.020000 --> 0:15:41.040000
 And it also has support for control characters,
 like new line, tab, carriage,

0:15:41.040000 --> 0:15:45.760000
 return, etc. So standard computing control
 characters, if you're an operating

0:15:45.760000 --> 0:15:51.600000
 system, a control character would be
 control or alt or caps lock, etc.

0:15:51.600000 --> 0:15:56.600000
 In terms of compatibility, ASCII is a subset
 of many other character encodings,

0:15:56.600000 --> 0:16:00.320000
 including much more comprehensive standards
 like Unicode, which is what

0:16:00.320000 --> 0:16:02.180000
 we'll take a look at shortly.

0:16:02.180000 --> 0:16:07.620000
 The first 128 characters of the Unicode
 standard correspond to the ASCII

0:16:07.620000 --> 0:16:13.000000
 characters. So Unicode actually has support
 for ASCII for backwards compatibility.

0:16:13.000000 --> 0:16:16.980000
 So in terms of its limitations, which
 is fairly obvious at this point,

0:16:16.980000 --> 0:16:21.000000
 ASCII is primarily designed for the
 English text and does not support

0:16:21.000000 --> 0:16:24.680000
 characters from other languages
 or special symbols.

0:16:24.680000 --> 0:16:26.900000
 That brings us to Unicode.

0:16:26.900000 --> 0:16:28.560000
 All right, so what is Unicode?

0:16:28.560000 --> 0:16:38.240000
 You probably are familiar with
 that's pretty much Unicode.

0:16:38.240000 --> 0:16:40.640000
 So Unicode is what is used today.

0:16:40.640000 --> 0:16:44.020000
 Unicode is a character set standard
 that aims to encompass characters

0:16:44.020000 --> 0:16:45.860000
 from all writing systems.

0:16:45.860000 --> 0:16:50.480000
 That's the keyword they're all writing
 systems and languages used worldwide.

0:16:50.480000 --> 0:16:54.180000
 So unlike early encoding standards
 like ASCII, which were limited to a

0:16:54.180000 --> 0:16:58.500000
 small set of characters, Unicode provides
 a unified system for representing

0:16:58.500000 --> 0:17:03.160000
 a vast range of characters, symbols
 and glyphs in a consistent manner.

0:17:03.160000 --> 0:17:07.900000
 So it now is pretty much a universal
 character set that has support for

0:17:07.900000 --> 0:17:11.360000
 all languages, right, or
 scripts, if you will.

0:17:11.360000 --> 0:17:14.880000
 And it enables computers to handle
 text and characters from a diverse

0:17:14.880000 --> 0:17:20.200000
 languages and scripts, making it essential
 for internationalization and

0:17:20.200000 --> 0:17:22.160000
 multilingual communication.

0:17:22.160000 --> 0:17:27.400000
 So again, we're not talking about the
 translation of different languages,

0:17:27.400000 --> 0:17:32.400000
 but just ensuring that different languages
 can be viewed or our different

0:17:32.400000 --> 0:17:35.020000
 languages can be represented correctly.

0:17:35.020000 --> 0:17:39.520000
 If there isn't a support for let's say
 Chinese characters, then your browser,

0:17:39.520000 --> 0:17:43.660000
 for example, in the case of the web
 will just display it in a very weird

0:17:43.660000 --> 0:17:45.340000
 collection of characters.

0:17:45.340000 --> 0:17:47.780000
 So it'll not be intelligible.

0:17:47.780000 --> 0:17:52.060000
 It'll not display the actual characters
 of the alphabet that, you know,

0:17:52.060000 --> 0:17:53.440000
 were intended to be displayed.

0:17:53.440000 --> 0:17:54.880000
 It'll just display weird symbols.

0:17:54.880000 --> 0:17:56.600000
 And I'll actually show you this.

0:17:56.600000 --> 0:17:59.020000
 So what does UTF stand for?

0:17:59.020000 --> 0:18:02.320000
 UTF stands for Unicode
 transformation format.

0:18:02.320000 --> 0:18:06.300000
 So it refers to a different character
 encoding scheme or different character

0:18:06.300000 --> 0:18:11.040000
 encoding schemes within the Unicode
 standard that are used to represent

0:18:11.040000 --> 0:18:13.960000
 Unicode characters as binary data.

0:18:13.960000 --> 0:18:18.420000
 All right, so that's where we have
 the concept of car of character set

0:18:18.420000 --> 0:18:20.220000
 encoding or char set encoding.

0:18:20.220000 --> 0:18:25.260000
 So character encoding is the representation
 in bytes of the symbols of

0:18:25.260000 --> 0:18:26.540000
 a character set.

0:18:26.540000 --> 0:18:29.600000
 Unicode has three main encoding systems.

0:18:29.600000 --> 0:18:33.580000
 We have UTF-8, which you're familiar
 with, which is widely used on the

0:18:33.580000 --> 0:18:37.040000
 web. We have UTF-16 and UTF-32.

0:18:37.040000 --> 0:18:40.920000
 So we'll really in this course only
 be touching on UTF-8, because again,

0:18:40.920000 --> 0:18:42.580000
 it makes the most sense.

0:18:42.580000 --> 0:18:47.140000
 And Unicode, as I said, stands for
 the Unicode transformation format.

0:18:47.140000 --> 0:18:49.960000
 So UTF-8, what does that mean?

0:18:49.960000 --> 0:18:53.920000
 Well, in terms of the abbreviation,
 it stands for Unicode transformation

0:18:53.920000 --> 0:18:59.620000
 eight bits. So UTF-8 is a variable
 length character encoding scheme.

0:18:59.620000 --> 0:19:04.520000
 And it utilizes eight bits or bytes,
 if you will, to represent characters.

0:19:04.520000 --> 0:19:07.560000
 ASCII characters are represented
 by a single byte.

0:19:07.560000 --> 0:19:09.960000
 That's for backwards compatibility.

0:19:09.960000 --> 0:19:14.340000
 And none ASCII characters are represented
 using multiple bytes with the

0:19:14.340000 --> 0:19:18.200000
 same number of bytes varying based
 on the character's code point.

0:19:18.200000 --> 0:19:22.940000
 UTF-8 is widely used on the web, and in
 many applications due to its efficiency

0:19:22.940000 --> 0:19:25.860000
 and compatibility with ASCII.

0:19:25.860000 --> 0:19:30.000000
 And then, you know, UTF-16, again, it
 just stands for Unicode transformation

0:19:30.000000 --> 0:19:35.360000
 format 16-bit. It is also a variable
 length character encoding scheme.

0:19:35.360000 --> 0:19:39.700000
 However, the difference is, as you know,
 it should already be aware, is

0:19:39.700000 --> 0:19:43.420000
 that it utilizes 16-bit units,
 which comprises of two bytes.

0:19:43.420000 --> 0:19:49.240000
 So two byte a byte is eight bits, eight
 bits times two 16 to represent

0:19:49.240000 --> 0:19:55.060000
 characters. And characters with
 code points below 65,536.

0:19:55.060000 --> 0:20:00.280000
 So that's the BMP basic multilingual plane
 are represented using two bytes.

0:20:00.280000 --> 0:20:04.760000
 And characters with higher code points
 that outside the BMP are represented

0:20:04.760000 --> 0:20:08.920000
 using four bytes, also known
 as surrogate pairs.

0:20:08.920000 --> 0:20:15.800000
 UTF-16 is commonly used in systems
 where that is required, as you can

0:20:15.800000 --> 0:20:19.800000
 obviously tell. So just a little bit
 more extensible with regards to the

0:20:19.800000 --> 0:20:24.220000
 range will not be touching on UTF-32, because
 it's sort of the same progression

0:20:24.220000 --> 0:20:28.540000
 in terms of support for a
 wider array of characters.

0:20:28.540000 --> 0:20:31.840000
 And UTF-8 is, as I said,
 utilized on the web.

0:20:31.840000 --> 0:20:33.900000
 So that's what we're going
 to be taking a look at.

0:20:33.900000 --> 0:20:37.780000
 So now I'm going to be putting everything
 that I've just explained into

0:20:37.780000 --> 0:20:41.760000
 context and giving you a practical example
 as to why this is important,

0:20:41.760000 --> 0:20:46.020000
 because you may be thinking to yourself,
 well, this is very, very complex.

0:20:46.020000 --> 0:20:49.200000
 How does all of this relate
 to web app pen testing?

0:20:49.200000 --> 0:20:50.340000
 Well, don't worry.

0:20:50.340000 --> 0:20:51.640000
 That's what I'm here to show you.

0:20:51.640000 --> 0:20:55.580000
 And it'll make even more sense as we
 approach the end of the encoding

0:20:55.580000 --> 0:20:59.980000
 section, because in the next video,
 for example, we'll be taking a look

0:20:59.980000 --> 0:21:04.400000
 at HTML encoding and then URL encoding
 and then base 64 encoding.

0:21:04.400000 --> 0:21:10.420000
 But you need to understand character
 set encoding in order for switch

0:21:10.420000 --> 0:21:12.320000
 over into my Kali Linux system.

0:21:12.320000 --> 0:21:14.700000
 There is no lab associated here.

0:21:14.700000 --> 0:21:18.020000
 But you can try this on your own again,
 even if you're using Windows,

0:21:18.020000 --> 0:21:20.680000
 it's a very simple demo
 you can go through.

0:21:20.680000 --> 0:21:22.740000
 But it'll highlight why
 it's so important.

0:21:22.740000 --> 0:21:26.440000
 So I'm going to switch over and we can
 then take a look at the example.

