Welcome!

Microsoft Cloud Authors: Janakiram MSV, Yeshim Deniz, David H Deans, Andreas Grabner, Stackify Blog

Related Topics: Microsoft Cloud

Microsoft Cloud: Article

Cover Story: Understanding Base64 Encoding

What it is, when to use it, and how to write custom Base64 encoding

If you work in a .NET environment you have probably come across Base64 encoded data. For example, Base64 encoding is used in ASP.NET for a Web application's ViewState value, as shown in Figure 1. Base64 encoding is also used to transmit binary data over e-mail. However, if you are like most of my colleagues (and me until recently) you do not have a thorough understanding of precisely what Base64 encoding is and when Base64 encoding should be used. In the this article I will explain exactly what Base64 encoding is, show you how to use the two primary .NET Framework methods that support Base64 encoding and decoding, and present a lightweight, custom C# implementation of Base64 encoding and decoding methods. This article assumes you are a .NET developer, tester, or manager and have intermediate level C# coding skill. After reading the article you'll have a solid grasp of Base64 encoding as well as the ability to write your own custom encoding methods. I think you'll find the ability to use Base64 encoded data is a valuable addition to your skill set.

The best way to show you where I'm headed in this article is with a screenshot. If you examine Figure 2 you'll see that I start with the arbitrary string "Hello" and use it to generate some binary data. After displaying the starting binary data in hexadecimal form, I convert the binary data to a Base64 encoded string using a method from the .NET Framework, and also using my custom encoding method. Notice the encoding results are the same. In the next part of the screenshot in Figure 2 I decode the Base64 encoded strings back to their original byte arrays, using both the built-in .NET Framework method and my custom implementation, and display in hexadecimal form. The complete program, which produced the screenshot in Figure 2, is presented in Listing 3.

What Is Base64 Encoding?
Exactly what is Base64 encoding? Base64 encoding is a scheme that encodes arbitrary binary data as a string composed from a set of 64 characters. The exact character set can be any 64 distinct ASCII characters, but by far the most common set is "A" through "Z," "a" through "z," "0" through "9," "+," and "/." For example, using this character set the 40-bit data:

01001000 01100101 01101100 01101100 01101111

can be Base64 encoded as the string:

SGVsbG8=

The trailing "=" character is a padding character as I will explain shortly. After seeing Base64 encoding for the first time, most engineers have an immediate question: Why would anyone want to use Base64 encoding? Base64 encoding is useful when you want to transmit binary data over a communication channel that is designed to transmit character data. For example, consider e-mail. The e-mail protocol SMTP was originally designed to send and receive only simple text data. However, suppose you want to transmit binary data such as a JPEG image. If you can encode the image as a Base64 string, then you can send the image just like any other message. Another common example is sending an ASP.NET Web application's ViewState value (which is a binary value representing the overall state of the application) over HTTP (which is an inherently text-based transport protocol). But why go to the trouble of Base64 encoding when "ordinary" encoding already exists? By ordinary encoding I mean regular hexadecimal encoding. For example, the 40-bit binary data above can be represented as a hexadecimal string: 48 65 6C 6C 6F (where the spaces are included just for readability). The answer is that Base64 encoding is more efficient than hexadecimal encoding in the sense that Base64 encoding requires fewer characters to represent the same data. Notice that the hexadecimal encoding of the 40-bit data above requires 10 characters while Base64 encoding requires only 8 characters - a 20 percent reduction.

Because Base64 encoding only uses 64 characters, any of the characters can be represented with just 6 bits because 26 = 64. Or, put another way, using 6 bits you can represent data in the range 000000b to 111111b, which is 0d to 63d. This is the key to Base64 encoding efficiency. The best way to explain how Base64 encoding works is with a picture as shown in Figure 3.

Suppose the first three bytes of input to be encoded are 48h, 65h, and 6Ch. In Figure 3 these values are shown in their binary representation: 01001000, 01100101, and 01101100. The first six bits of input, 010010, have value 18d, which in turn maps to character [18] in the Base64 character set, which is "S." The second six bits of input - the last two bits of the first byte of input and the first four bits of the second byte of input - equal 6d, which maps to character "G," and so on. Notice that three bytes of input map neatly to four character of output. Because of this it is convenient to implement Base64 encoding in "blocks" that represent a group of three bytes of input, or four characters of output.

NET Support for Base64 Encoding
The .NET Framework supports Base64 encoding with two methods. The Convert.ToBase64String() method accepts a byte array as an input argument and returns a Base64 encoded string (using the usual 64-character set described in the previous section). The Convert.FromBase64String() accepts a string argument (which is assumed to be a valid Base64 encoded string) and returns the corresponding byte array. Using these two methods is very easy. Notice both methods belong to the Convert class and are static so you do not need to instantiate an object to use the methods. The Convert class is part of the System namespace. Consider this code snippet:

byte[] bytes = new byte[] { 0x5F, 0xC9, 0xBF, 0x17 };
string base64 = Convert.ToBase64String(bytes);
Console.WriteLine(base64);

This code would produce "X8m/Fw==" as output. Notice that the input has size 4 bytes. The first three bytes (3 * 8 = 24 bits) of input are used to produce the first four Base64 characters (4 * 6 = 24 bits) of output. The last input byte produces the rest of the output, and the output is padded with "=" characters to bring the output size up to an even multiple of 4. The technique to decode is similar. The statements

string encoded = "X8m/Fw==";
byte[] result = Convert.FromBase64String(encoded);
Console.WriteLine(BitConverter.ToString(result));

produce "5F-C9-BF-17" as output. The Framework Base64 methods are simple and straightforward. However, if you are encoding, transmitting, and decoding large amounts of binary data, it is up to you write auxiliary code, which buffers the process by breaking the input data into manageable-sized chunks.

A Lightweight Custom Base64 Encoder
The .NET Framework's ToBase64-String() and FromBase64String() methods will meet most of your Base64 encoding needs. However what if you are developing a system and need a slightly different encoding scheme? For instance, you may want to use a different character set than the normal{"A"-"Z," "a"-"z," "0"-9," "+," "/"} set. In this section I'll present a lightweight, custom Base64 encoder written in C# that you can use as a starting point for your own custom encoder. If you search the Internet you'll find quite a few Base64 encoding examples. The one I present here is a hybrid of several I found combined with one I wrote recently, and is designed for maximum clarity rather than for efficiency. The custom encoder is presented in Listing 1.

Because I want my customer encoder to mimic the interface of the Framework encoder, I begin by creating an overall structure of:

public class MyConverter
{
    public static string ToBase64String(byte[] value)
   {
      // implementation goes here
}
} // class MyConverter

Of course there are many other design alternatives, but making the custom encoder signature the same as the Framework's encoder signature makes sense. I begin my encoder implementation by declaring an array of the 64 characters I want to use for my encoding:

char[] base64Chars = new char[]
{ 'A','B','C','D','E','F','G','H','I','J','K','L','M',
'N','O','P','Q','R','S','T','U','V','W','X','Y','Z',
'a','b','c','d','e','f','g','h','i','j','k','l','m',
'n','o','p','q','r','s','t','u','v','w','x','y','z',
'0','1','2','3','4','5','6','7','8','9','+','/' };

This array acts as a lookup table to map a decimal value in the range 0 - 63 to a Base64 character. For example, 0 maps to "A," 1 maps to "B," and 26 maps to "a." I use the normal character set but you can use different characters, or change the order for a custom encoding scheme. Next I compute two values that will control the encoding algorithm:

int numBlocks;
int padBytes;
if ((value.Length % 3) == 0)
{
     numBlocks = value.Length / 3;
     padBytes = 0;
}
else
{
     numBlocks = 1 + (value.Length / 3);
     padBytes = 3 - (value.Length % 3);
}

More Stories By James McCaffrey

Dr. James McCaffrey works for Volt Information Sciences, Inc., where he manages technical training for software engineers working at Microsoft's Redmond, WA campus. He has worked on several Microsoft products, including Internet Explorer and MSN Search. James can be reached at [email protected] or [email protected]

Comments (3) View Comments

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


Most Recent Comments
trash_incinerator 02/19/10 01:17:00 PM EST

With all due respect to the author, your explanation for why Base64 encoding exists is wrong. The "string" of hex for your example is not comprised of 10 characters as you have indicated. Hex uses 2 digits to represent 8 bits. Four bits for the first digit and four bits for the second. There are, in fact, only 5 bytes of data there, whereas the Base64 encoded string is using 8 bytes.

Base64 is NOT a means of compressing data. In fact it makes the information being represented larger. The reason why this is sometimes necessary is because of the fact that systems assign special meanings to specific bytes or byte sequences. In XML for example, there are special bytes that are not considered valid characters in XML. In order to send information in an XML stream with characters that are not allowed, you have to replace the illegal bytes with legal ones. Hence, Base64 allows you to take arbitrary bytes and reassign them in a way that can be reversed later.

Kumanan Murugesan 04/16/08 10:07:55 AM EDT

Dr. James,
Wonderful article. I was wondering what this does and why is it required many times like other folks.

SYS-CON Belgium News Desk 03/19/06 10:04:38 AM EST

If you work in a .NET environment you have probably come across Base64 encoded data. For example, Base64 encoding is used in ASP.NET for a Web application's ViewState value, as shown in Figure 1. Base64 encoding is also used to transmit binary data over e-mail. However, if you are like most of my colleagues (and me until recently) you do not have a thorough understanding of precisely what Base64 encoding is and when Base64 encoding should be used. In the this article I will explain exactly what Base64 encoding is, show you how to use the two primary .NET Framework methods that support Base64 encoding and decoding, and present a lightweight, custom C# implementation of Base64 encoding and decoding methods. This article assumes you are a .NET developer, tester, or manager and have intermediate level C# coding skill. After reading the article you'll have a solid grasp of Base64 encoding as well as the ability to write your own custom encoding methods. I think you'll find the ability to use Base64 encoded data is a valuable addition to your skill set.

@ThingsExpo Stories
Agile has finally jumped the technology shark, expanding outside the software world. Enterprises are now increasingly adopting Agile practices across their organizations in order to successfully navigate the disruptive waters that threaten to drown them. In our quest for establishing change as a core competency in our organizations, this business-centric notion of Agile is an essential component of Agile Digital Transformation. In the years since the publication of the Agile Manifesto, the conn...
SYS-CON Events announced today that App2Cloud will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct. 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. App2Cloud is an online Platform, specializing in migrating legacy applications to any Cloud Providers (AWS, Azure, Google Cloud).
WebRTC is great technology to build your own communication tools. It will be even more exciting experience it with advanced devices, such as a 360 Camera, 360 microphone, and a depth sensor camera. In his session at @ThingsExpo, Masashi Ganeko, a manager at INFOCOM Corporation, will introduce two experimental projects from his team and what they learned from them. "Shotoku Tamago" uses the robot audition software HARK to track speakers in 360 video of a remote party. "Virtual Teleport" uses a mu...
Internet of @ThingsExpo, taking place October 31 - November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 21st Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The Internet of Things (IoT) is the most profound change in personal and enterprise IT since the creation of the Worldwide Web more than 20 years ago. All major researchers estimate there will be tens of billions devic...
Mobile device usage has increased exponentially during the past several years, as consumers rely on handhelds for everything from news and weather to banking and purchases. What can we expect in the next few years? The way in which we interact with our devices will fundamentally change, as businesses leverage Artificial Intelligence. We already see this taking shape as businesses leverage AI for cost savings and customer responsiveness. This trend will continue, as AI is used for more sophistica...
SYS-CON Events announced today that SourceForge has been named “Media Sponsor” of SYS-CON's 21st International Cloud Expo, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. SourceForge is the largest, most trusted destination for Open Source Software development, collaboration, discovery and download on the web serving over 32 million viewers, 150 million downloads and over 460,000 active development projects each and every month.
SYS-CON Events announced today that Massive Networks will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Massive Networks mission is simple. To help your business operate seamlessly with fast, reliable, and secure internet and network solutions. Improve your customer's experience with outstanding connections to your cloud.
SYS-CON Events announced today that DXWorldExpo has been named “Global Sponsor” of SYS-CON's 21st International Cloud Expo, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Digital Transformation is the key issue driving the global enterprise IT business. Digital Transformation is most prominent among Global 2000 enterprises and government institutions.
SYS-CON Events announced today that WineSOFT will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Based in Seoul and Irvine, WineSOFT is an innovative software house focusing on internet infrastructure solutions. The venture started as a bootstrap start-up in 2010 by focusing on making the internet faster and more powerful. WineSOFT’s knowledge is based on the expertise of TCP/IP, VPN, SS...
SYS-CON Events announced today that Akvelon will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Akvelon is a business and technology consulting firm that specializes in applying cutting-edge technology to problems in fields as diverse as mobile technology, sports technology, finance, and healthcare.
SYS-CON Events announced today that TechTarget has been named “Media Sponsor” of SYS-CON's 21st International Cloud Expo, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. TechTarget storage websites are the best online information resource for news, tips and expert advice for the storage, backup and disaster recovery markets.
Real IoT production deployments running at scale are collecting sensor data from hundreds / thousands / millions of devices. The goal is to take business-critical actions on the real-time data and find insights from stored datasets. In his session at @ThingsExpo, John Walicki, Watson IoT Developer Advocate at IBM Cloud, will provide a fast-paced developer journey that follows the IoT sensor data from generation, to edge gateway, to edge analytics, to encryption, to the IBM Bluemix cloud, to Wa...
SYS-CON Events announced today that Dasher Technologies will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Dasher Technologies, Inc. ® is a premier IT solution provider that delivers expert technical resources along with trusted account executives to architect and deliver complete IT solutions and services to help our clients execute their goals, plans and objectives. Since 1999, we've...
DevOps at Cloud Expo – being held October 31 - November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA – announces that its Call for Papers is open. Born out of proven success in agile development, cloud computing, and process automation, DevOps is a macro trend you cannot afford to miss. From showcase success stories from early adopters and web-scale businesses, DevOps is expanding to organizations of all sizes, including the world's largest enterprises – and delivering real r...
No hype cycles or predictions of a gazillion things here. IoT is here. You get it. You know your business and have great ideas for a business transformation strategy. What comes next? Time to make it happen. In his session at @ThingsExpo, Jay Mason, an Associate Partner of Analytics, IoT & Cybersecurity at M&S Consulting, will present a step-by-step plan to develop your technology implementation strategy. He will discuss the evaluation of communication standards and IoT messaging protocols, dat...
SYS-CON Events announced today that Massive Networks, that helps your business operate seamlessly with fast, reliable, and secure internet and network solutions, has been named "Exhibitor" of SYS-CON's 21st International Cloud Expo ®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. As a premier telecommunications provider, Massive Networks is headquartered out of Louisville, Colorado. With years of experience under their belt, their team of...
Elon Musk is among the notable industry figures who worries about the power of AI to destroy rather than help society. Mark Zuckerberg, on the other hand, embraces all that is going on. AI is most powerful when deployed across the vast networks being built for Internets of Things in the manufacturing, transportation and logistics, retail, healthcare, government and other sectors. Is AI transforming IoT for the good or the bad? Do we need to worry about its potential destructive power? Or will we...
SYS-CON Events announced today that App2Cloud will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct. 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. App2Cloud is an online Platform, specializing in migrating legacy applications to any Cloud Providers (AWS, Azure, Google Cloud).
SYS-CON Events announced today that MobiDev, a client-oriented software development company, will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place October 31-November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. MobiDev is a software company that develops and delivers turn-key mobile apps, websites, web services, and complex software systems for startups and enterprises. Since 2009 it has grown from a small group of passionate engineers and business...
With major technology companies and startups seriously embracing Cloud strategies, now is the perfect time to attend 21st Cloud Expo October 31 - November 2, 2017, at the Santa Clara Convention Center, CA, and June 12-14, 2018, at the Javits Center in New York City, NY, and learn what is going on, contribute to the discussions, and ensure that your enterprise is on the right path to Digital Transformation.