Welcome!

Microsoft Cloud Authors: Liz McMillan, David H Deans, Automic Blog, Pat Romanski, Janakiram MSV

Related Topics: Microsoft Cloud

Microsoft Cloud: Article

Cover Story: Understanding Base64 Encoding

What it is, when to use it, and how to write custom Base64 encoding

If you work in a .NET environment you have probably come across Base64 encoded data. For example, Base64 encoding is used in ASP.NET for a Web application's ViewState value, as shown in Figure 1. Base64 encoding is also used to transmit binary data over e-mail. However, if you are like most of my colleagues (and me until recently) you do not have a thorough understanding of precisely what Base64 encoding is and when Base64 encoding should be used. In the this article I will explain exactly what Base64 encoding is, show you how to use the two primary .NET Framework methods that support Base64 encoding and decoding, and present a lightweight, custom C# implementation of Base64 encoding and decoding methods. This article assumes you are a .NET developer, tester, or manager and have intermediate level C# coding skill. After reading the article you'll have a solid grasp of Base64 encoding as well as the ability to write your own custom encoding methods. I think you'll find the ability to use Base64 encoded data is a valuable addition to your skill set.

The best way to show you where I'm headed in this article is with a screenshot. If you examine Figure 2 you'll see that I start with the arbitrary string "Hello" and use it to generate some binary data. After displaying the starting binary data in hexadecimal form, I convert the binary data to a Base64 encoded string using a method from the .NET Framework, and also using my custom encoding method. Notice the encoding results are the same. In the next part of the screenshot in Figure 2 I decode the Base64 encoded strings back to their original byte arrays, using both the built-in .NET Framework method and my custom implementation, and display in hexadecimal form. The complete program, which produced the screenshot in Figure 2, is presented in Listing 3.

What Is Base64 Encoding?
Exactly what is Base64 encoding? Base64 encoding is a scheme that encodes arbitrary binary data as a string composed from a set of 64 characters. The exact character set can be any 64 distinct ASCII characters, but by far the most common set is "A" through "Z," "a" through "z," "0" through "9," "+," and "/." For example, using this character set the 40-bit data:

01001000 01100101 01101100 01101100 01101111

can be Base64 encoded as the string:

SGVsbG8=

The trailing "=" character is a padding character as I will explain shortly. After seeing Base64 encoding for the first time, most engineers have an immediate question: Why would anyone want to use Base64 encoding? Base64 encoding is useful when you want to transmit binary data over a communication channel that is designed to transmit character data. For example, consider e-mail. The e-mail protocol SMTP was originally designed to send and receive only simple text data. However, suppose you want to transmit binary data such as a JPEG image. If you can encode the image as a Base64 string, then you can send the image just like any other message. Another common example is sending an ASP.NET Web application's ViewState value (which is a binary value representing the overall state of the application) over HTTP (which is an inherently text-based transport protocol). But why go to the trouble of Base64 encoding when "ordinary" encoding already exists? By ordinary encoding I mean regular hexadecimal encoding. For example, the 40-bit binary data above can be represented as a hexadecimal string: 48 65 6C 6C 6F (where the spaces are included just for readability). The answer is that Base64 encoding is more efficient than hexadecimal encoding in the sense that Base64 encoding requires fewer characters to represent the same data. Notice that the hexadecimal encoding of the 40-bit data above requires 10 characters while Base64 encoding requires only 8 characters - a 20 percent reduction.

Because Base64 encoding only uses 64 characters, any of the characters can be represented with just 6 bits because 26 = 64. Or, put another way, using 6 bits you can represent data in the range 000000b to 111111b, which is 0d to 63d. This is the key to Base64 encoding efficiency. The best way to explain how Base64 encoding works is with a picture as shown in Figure 3.

Suppose the first three bytes of input to be encoded are 48h, 65h, and 6Ch. In Figure 3 these values are shown in their binary representation: 01001000, 01100101, and 01101100. The first six bits of input, 010010, have value 18d, which in turn maps to character [18] in the Base64 character set, which is "S." The second six bits of input - the last two bits of the first byte of input and the first four bits of the second byte of input - equal 6d, which maps to character "G," and so on. Notice that three bytes of input map neatly to four character of output. Because of this it is convenient to implement Base64 encoding in "blocks" that represent a group of three bytes of input, or four characters of output.

NET Support for Base64 Encoding
The .NET Framework supports Base64 encoding with two methods. The Convert.ToBase64String() method accepts a byte array as an input argument and returns a Base64 encoded string (using the usual 64-character set described in the previous section). The Convert.FromBase64String() accepts a string argument (which is assumed to be a valid Base64 encoded string) and returns the corresponding byte array. Using these two methods is very easy. Notice both methods belong to the Convert class and are static so you do not need to instantiate an object to use the methods. The Convert class is part of the System namespace. Consider this code snippet:

byte[] bytes = new byte[] { 0x5F, 0xC9, 0xBF, 0x17 };
string base64 = Convert.ToBase64String(bytes);
Console.WriteLine(base64);

This code would produce "X8m/Fw==" as output. Notice that the input has size 4 bytes. The first three bytes (3 * 8 = 24 bits) of input are used to produce the first four Base64 characters (4 * 6 = 24 bits) of output. The last input byte produces the rest of the output, and the output is padded with "=" characters to bring the output size up to an even multiple of 4. The technique to decode is similar. The statements

string encoded = "X8m/Fw==";
byte[] result = Convert.FromBase64String(encoded);
Console.WriteLine(BitConverter.ToString(result));

produce "5F-C9-BF-17" as output. The Framework Base64 methods are simple and straightforward. However, if you are encoding, transmitting, and decoding large amounts of binary data, it is up to you write auxiliary code, which buffers the process by breaking the input data into manageable-sized chunks.

A Lightweight Custom Base64 Encoder
The .NET Framework's ToBase64-String() and FromBase64String() methods will meet most of your Base64 encoding needs. However what if you are developing a system and need a slightly different encoding scheme? For instance, you may want to use a different character set than the normal{"A"-"Z," "a"-"z," "0"-9," "+," "/"} set. In this section I'll present a lightweight, custom Base64 encoder written in C# that you can use as a starting point for your own custom encoder. If you search the Internet you'll find quite a few Base64 encoding examples. The one I present here is a hybrid of several I found combined with one I wrote recently, and is designed for maximum clarity rather than for efficiency. The custom encoder is presented in Listing 1.

Because I want my customer encoder to mimic the interface of the Framework encoder, I begin by creating an overall structure of:

public class MyConverter
{
    public static string ToBase64String(byte[] value)
   {
      // implementation goes here
}
} // class MyConverter

Of course there are many other design alternatives, but making the custom encoder signature the same as the Framework's encoder signature makes sense. I begin my encoder implementation by declaring an array of the 64 characters I want to use for my encoding:

char[] base64Chars = new char[]
{ 'A','B','C','D','E','F','G','H','I','J','K','L','M',
'N','O','P','Q','R','S','T','U','V','W','X','Y','Z',
'a','b','c','d','e','f','g','h','i','j','k','l','m',
'n','o','p','q','r','s','t','u','v','w','x','y','z',
'0','1','2','3','4','5','6','7','8','9','+','/' };

This array acts as a lookup table to map a decimal value in the range 0 - 63 to a Base64 character. For example, 0 maps to "A," 1 maps to "B," and 26 maps to "a." I use the normal character set but you can use different characters, or change the order for a custom encoding scheme. Next I compute two values that will control the encoding algorithm:

int numBlocks;
int padBytes;
if ((value.Length % 3) == 0)
{
     numBlocks = value.Length / 3;
     padBytes = 0;
}
else
{
     numBlocks = 1 + (value.Length / 3);
     padBytes = 3 - (value.Length % 3);
}

More Stories By James McCaffrey

Dr. James McCaffrey works for Volt Information Sciences, Inc., where he manages technical training for software engineers working at Microsoft's Redmond, WA campus. He has worked on several Microsoft products, including Internet Explorer and MSN Search. James can be reached at [email protected] or [email protected]

Comments (3) View Comments

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


Most Recent Comments
trash_incinerator 02/19/10 01:17:00 PM EST

With all due respect to the author, your explanation for why Base64 encoding exists is wrong. The "string" of hex for your example is not comprised of 10 characters as you have indicated. Hex uses 2 digits to represent 8 bits. Four bits for the first digit and four bits for the second. There are, in fact, only 5 bytes of data there, whereas the Base64 encoded string is using 8 bytes.

Base64 is NOT a means of compressing data. In fact it makes the information being represented larger. The reason why this is sometimes necessary is because of the fact that systems assign special meanings to specific bytes or byte sequences. In XML for example, there are special bytes that are not considered valid characters in XML. In order to send information in an XML stream with characters that are not allowed, you have to replace the illegal bytes with legal ones. Hence, Base64 allows you to take arbitrary bytes and reassign them in a way that can be reversed later.

Kumanan Murugesan 04/16/08 10:07:55 AM EDT

Dr. James,
Wonderful article. I was wondering what this does and why is it required many times like other folks.

SYS-CON Belgium News Desk 03/19/06 10:04:38 AM EST

If you work in a .NET environment you have probably come across Base64 encoded data. For example, Base64 encoding is used in ASP.NET for a Web application's ViewState value, as shown in Figure 1. Base64 encoding is also used to transmit binary data over e-mail. However, if you are like most of my colleagues (and me until recently) you do not have a thorough understanding of precisely what Base64 encoding is and when Base64 encoding should be used. In the this article I will explain exactly what Base64 encoding is, show you how to use the two primary .NET Framework methods that support Base64 encoding and decoding, and present a lightweight, custom C# implementation of Base64 encoding and decoding methods. This article assumes you are a .NET developer, tester, or manager and have intermediate level C# coding skill. After reading the article you'll have a solid grasp of Base64 encoding as well as the ability to write your own custom encoding methods. I think you'll find the ability to use Base64 encoded data is a valuable addition to your skill set.

@ThingsExpo Stories
SYS-CON Events announced today that SoftLayer, an IBM Company, has been named “Gold Sponsor” of SYS-CON's 18th Cloud Expo, which will take place on June 7-9, 2016, at the Javits Center in New York, New York. SoftLayer, an IBM Company, provides cloud infrastructure as a service from a growing number of data centers and network points of presence around the world. SoftLayer’s customers range from Web startups to global enterprises.
SYS-CON Events announced today that Auditwerx will exhibit at SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. Auditwerx specializes in SOC 1, SOC 2, and SOC 3 attestation services throughout the U.S. and Canada. As a division of Carr, Riggs & Ingram (CRI), one of the top 20 largest CPA firms nationally, you can expect the resources, skills, and experience of a much larger firm combined with the accessibility and attent...
SYS-CON Events announced today that CA Technologies has been named “Platinum Sponsor” of SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY, and the 21st International Cloud Expo®, which will take place October 31-November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. CA Technologies helps customers succeed in a future where every business – from apparel to energy – is being rewritten by software. From ...
SYS-CON Events announced today that Technologic Systems Inc., an embedded systems solutions company, will exhibit at SYS-CON's @ThingsExpo, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. Technologic Systems is an embedded systems company with headquarters in Fountain Hills, Arizona. They have been in business for 32 years, helping more than 8,000 OEM customers and building over a hundred COTS products that have never been discontinued. Technologic Systems’ pr...
SYS-CON Events announced today that HTBase will exhibit at SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. HTBase (Gartner 2016 Cool Vendor) delivers a Composable IT infrastructure solution architected for agility and increased efficiency. It turns compute, storage, and fabric into fluid pools of resources that are easily composed and re-composed to meet each application’s needs. With HTBase, companies can quickly prov...
SYS-CON Events announced today that Loom Systems will exhibit at SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. Founded in 2015, Loom Systems delivers an advanced AI solution to predict and prevent problems in the digital business. Loom stands alone in the industry as an AI analysis platform requiring no prior math knowledge from operators, leveraging the existing staff to succeed in the digital era. With offices in S...
Buzzword alert: Microservices and IoT at a DevOps conference? What could possibly go wrong? In this Power Panel at DevOps Summit, moderated by Jason Bloomberg, the leading expert on architecting agility for the enterprise and president of Intellyx, panelists peeled away the buzz and discuss the important architectural principles behind implementing IoT solutions for the enterprise. As remote IoT devices and sensors become increasingly intelligent, they become part of our distributed cloud enviro...
SYS-CON Events announced today that T-Mobile will exhibit at SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. As America's Un-carrier, T-Mobile US, Inc., is redefining the way consumers and businesses buy wireless services through leading product and service innovation. The Company's advanced nationwide 4G LTE network delivers outstanding wireless experiences to 67.4 million customers who are unwilling to compromise on ...
SYS-CON Events announced today that Infranics will exhibit at SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. Since 2000, Infranics has developed SysMaster Suite, which is required for the stable and efficient management of ICT infrastructure. The ICT management solution developed and provided by Infranics continues to add intelligence to the ICT infrastructure through the IMC (Infra Management Cycle) based on mathemat...
SYS-CON Events announced today that Interoute, owner-operator of one of Europe's largest networks and a global cloud services platform, has been named “Bronze Sponsor” of SYS-CON's 20th Cloud Expo, which will take place on June 6-8, 2017 at the Javits Center in New York, New York. Interoute is the owner-operator of one of Europe's largest networks and a global cloud services platform which encompasses 12 data centers, 14 virtual data centers and 31 colocation centers, with connections to 195 add...
SYS-CON Events announced today that Cloudistics, an on-premises cloud computing company, has been named “Bronze Sponsor” of SYS-CON's 20th International Cloud Expo®, which will take place on June 6-8, 2017, at the Javits Center in New York City, NY. Cloudistics delivers a complete public cloud experience with composable on-premises infrastructures to medium and large enterprises. Its software-defined technology natively converges network, storage, compute, virtualization, and management into a ...
In his session at @ThingsExpo, Eric Lachapelle, CEO of the Professional Evaluation and Certification Board (PECB), will provide an overview of various initiatives to certifiy the security of connected devices and future trends in ensuring public trust of IoT. Eric Lachapelle is the Chief Executive Officer of the Professional Evaluation and Certification Board (PECB), an international certification body. His role is to help companies and individuals to achieve professional, accredited and worldw...
In his General Session at 16th Cloud Expo, David Shacochis, host of The Hybrid IT Files podcast and Vice President at CenturyLink, investigated three key trends of the “gigabit economy" though the story of a Fortune 500 communications company in transformation. Narrating how multi-modal hybrid IT, service automation, and agile delivery all intersect, he will cover the role of storytelling and empathy in achieving strategic alignment between the enterprise and its information technology.
Microservices are a very exciting architectural approach that many organizations are looking to as a way to accelerate innovation. Microservices promise to allow teams to move away from monolithic "ball of mud" systems, but the reality is that, in the vast majority of organizations, different projects and technologies will continue to be developed at different speeds. How to handle the dependencies between these disparate systems with different iteration cycles? Consider the "canoncial problem" ...
The Internet of Things is clearly many things: data collection and analytics, wearables, Smart Grids and Smart Cities, the Industrial Internet, and more. Cool platforms like Arduino, Raspberry Pi, Intel's Galileo and Edison, and a diverse world of sensors are making the IoT a great toy box for developers in all these areas. In this Power Panel at @ThingsExpo, moderated by Conference Chair Roger Strukhoff, panelists discussed what things are the most important, which will have the most profound e...
Keeping pace with advancements in software delivery processes and tooling is taxing even for the most proficient organizations. Point tools, platforms, open source and the increasing adoption of private and public cloud services requires strong engineering rigor - all in the face of developer demands to use the tools of choice. As Agile has settled in as a mainstream practice, now DevOps has emerged as the next wave to improve software delivery speed and output. To make DevOps work, organization...
My team embarked on building a data lake for our sales and marketing data to better understand customer journeys. This required building a hybrid data pipeline to connect our cloud CRM with the new Hadoop Data Lake. One challenge is that IT was not in a position to provide support until we proved value and marketing did not have the experience, so we embarked on the journey ourselves within the product marketing team for our line of business within Progress. In his session at @BigDataExpo, Sum...
Web Real-Time Communication APIs have quickly revolutionized what browsers are capable of. In addition to video and audio streams, we can now bi-directionally send arbitrary data over WebRTC's PeerConnection Data Channels. With the advent of Progressive Web Apps and new hardware APIs such as WebBluetooh and WebUSB, we can finally enable users to stitch together the Internet of Things directly from their browsers while communicating privately and securely in a decentralized way.
DevOps is often described as a combination of technology and culture. Without both, DevOps isn't complete. However, applying the culture to outdated technology is a recipe for disaster; as response times grow and connections between teams are delayed by technology, the culture will die. A Nutanix Enterprise Cloud has many benefits that provide the needed base for a true DevOps paradigm.
What sort of WebRTC based applications can we expect to see over the next year and beyond? One way to predict development trends is to see what sorts of applications startups are building. In his session at @ThingsExpo, Arin Sime, founder of WebRTC.ventures, will discuss the current and likely future trends in WebRTC application development based on real requests for custom applications from real customers, as well as other public sources of information,