Welcome!

Microsoft Cloud Authors: Janakiram MSV, Pat Romanski, Steven Mandel, John Basso, Liz McMillan

Related Topics: Microsoft Cloud

Microsoft Cloud: Article

WinFS: The Windows Storage Platform of the Future

How to Write Applications that Take Advantage of WinFS

The next version of the Windows operating system (codenamed "Longhorn") has a new storage subsystem (codenamed "WinFS"). In this article, we will try to understand the need for "WinFS"; define "WinFS", its type system, and its data model; and learn how to write applications that take advantage of "WinFS".

Why Does Windows Need a New Storage Subsystem?
The hardware industry is confidently striding towards conquering the 3T challenge (a teraflops processor, a terabyte hard disk, and a terabits per sec bandwidth). Increase in the size of hard disk storage has been complemented by the exponential increase in the production of digital data. The amount of digital data being born daily is so phenomenal that the pen, the paper, and the typewriter have achieved the status of endangered stationery. Digital data is stored by operating systems on magnetic media - like hard disks. Windows uses file systems like NTFS and FAT32 to organize data on hard disks that are divided into logical drives. Each drive has a root folder, which can have one or more files and folders. Data is stored in files.

Data has personality. It can be documents (text, DOC, PDF, RTF, PostScript formats), images (bitmap, GIF, JPEG, PNG, WMF formats), audio (WAV, MP3, WMA, AIFF, formats), video (MOV, MPEG, AVI, ASF, QuickTime formats), and more. Along with data, these formats also store rich metadata. NTFS and FAT32 are capable of reading only basic metadata from files, like their name, size, date modified, and type (derived from the file extension). So, if I want to search all songs of a particular artist, I can't do it just by using the file system. I'll need a specialized application like Windows Media Player, which is capable of reading the rich metadata provided by all audio files.

What is "WinFS"?
"WinFS" is the Windows storage platform of the future. It solves many of the problems associated with current file systems by storing data and its metadata together. This way data can be organized, searched, shared, and related depending upon what it is and not what it's called (filename)! "WinFS" also makes it easy to deal with non-file-backed data like personal contacts, e-mail messages, and event calendar.

"WinFS" allows a user to organize data with great flexibility. Using "WinFS", a user can group data according to common characteristics, create a customized containment hierarchy to hold data, and associate one piece of data to another via relationships. This organizational flexibility provides the user with greater flexibility when searching for data. The user can search for data based on its attributes, its relationship with other data, and by its storage location.

Searching Data
Because of its ability to organize data based on various parameters, "WinFS" presents data as a directed acyclic graph (DAG) instead of as a tree (as in NTFS and FAT32). In NTFS and FAT32, since data is stored as a tree, it can be located based only on a single criteria (its directory path), while in "WinFS", since data is stored as a graph it can be searched based on multiple criteria.

Sharing Data
"WinFS" facilitates data sharing among various applications by giving file-backed data (documents, audio video, media, and so on) as well as non file-backed (e-mail, contacts, appointments, and so on), permanent residency status in the operating system and providing a unique API to access them. Now different applications can use these common pieces of data without worrying about maintaining individual data stores and the associated synchronization.

WinFS Building Blocks
Figure 1 illustrates the various building blocks of "WinFS".

Core "WinFS"
"WinFS" does not supplant NTFS; rather, it utilizes the basic file system services of NTFS. The core building block of "WinFS" includes the relational storage engine that provides file system services like ACL support, import/export, and quota management.

Data Model
"WinFS" defines a rich data model that resides on top of a relational storage engine. "WinFS" represents a piece of data (item) as a tuple in a relation. The attributes of the tuple describe the piece of data. Items can be related to each other by defining relationships between the tuples. "WinFS" also provides the ability to extend items and relationships.

Schemas
"WinFS" has built-in schemas to understand the rich metadata associated with your data. Some of the built-in schemas are for common data like documents, e-mail, contacts, appointments, tasks, and more. "WinFS" also allows you to write custom schemas for your own data.

"WinFS" provides certain services like synchronization and rules. These services are layered on top of core "WinFS" and provide "WinFS" with extended functionality. For example, the synchronization service enables you to synchronize two or more "WinFS" stores.

APIs
"WinFS"provides support for multiple programming models: object-oriented, relational, and XML based. "WinFS" can also be programmed using the Win32 API.

"WinFS" Type System Basics
"WinFS" is a "strongly typed" storage system. All data stored in "WinFS" is typed; that is, it is an object of some "WinFS" type. To understand the "WinFS" type system, we first have to understand the following four concepts:

  • Items: All data is stored in WinFS  as a type specialized from Item (System.Storage.Item). Examples of types that derive from the Item are: Contact, Document, Task, and Event (all of which are found in the System.Storage.Core namespace)
  • ScalarTypes: An atomic piece of information that describes an item.
  • NestedTypes: A set of information that can be stored about an item. A nested type can have other nested types and scalar types in it.
  • Relationships: Relate one item (source) to another (target). Relationships can be of three types:
    - Holding relationship: The source controls the lifetime of the target. These relationships are many-to-many, that is, a source can "hold" multiple targets and a target can be "held" by multiple sources. If a target looses all sources, it's deleted.
    - Reference relationship: Similar to a holding relationship but without the lifetime management of the target.
    - Embedding relationship: The source embeds the target. Strictly one-to-one.

Programming "WinFS"

In this section, we'll take a brief look at the object-oriented managed API provided by "WinFS". This API allows us to search, relate, and act upon data stored by WinFS

.An installation of Longhorn will have a single instance of "WinFS" service running on it. A "WinFS" instance can maintain multiple data stores. Each data source is referenced by its UNC path syntax as shown below:

\\<machine-name\<data-store-name>

All "WinFS" instances have a default store (DefaultStore). The default store on your "Longhorn" system will be called:

\\localhost\defaultstore

To program "WinFS", you have to get hold of an ItemContext object. This is done by calling the static Open method of the ItemContext class. The method is passed the UNC path of the store you want to program. If you pass nothing, the ItemContext refers to the default store.

All the types that are required to program "WinFS" can be found in three assemblies: System.Storage. dll, System.Storage.Schema.dll, and WinCorLib.dll.

As an example, Listing 1 is code that prints out all folders in the default store.

A very useful class provided by "WinFS" is the ItemSearcher (System.Storage.ItemSearcher) class, which allows you to search for items in a particular item type. Each item type (for example, Contact, Document, Folder, Event, and so on) has a static method called GetSearcher that takes in the ItemContext object and returns the ItemSearcher. Using the ItemSearcher, you can search for the items you need.

For example, Listing 2 uses the ItemSearcher to search for all documents whose title begins with the string 'Result'.

Conclusion
In this brief survey of "WinFS", we have seen that the future of data is with context-based storage systems. "WinFS" will provide the basis on which advanced data-mining tools are going to be built for the personal computer.

More Stories By Mujtaba Syed

Mujtaba Syed works as a software architect with Marlabs Inc. He is an MCSD
(early achiever) and loves to speak about and write on Microsoft .NET. Mujtaba has been programming the Microsoft .NET Framework since its beta 1 release. His current interests are focused on Longhorn.

Comments (2) View Comments

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


Most Recent Comments
Mujtaba Syed 06/08/04 03:11:22 PM EDT

It''s corrected now.

Mujtaba Syed 06/08/04 02:29:18 PM EDT

Somehow the online version is showing WinFX at all places it should show WinFS! We will get this rectified ASAP. Thanks.

@ThingsExpo Stories
Why do your mobile transformations need to happen today? Mobile is the strategy that enterprise transformation centers on to drive customer engagement. In his general session at @ThingsExpo, Roger Woods, Director, Mobile Product & Strategy – Adobe Marketing Cloud, covered key IoT and mobile trends that are forcing mobile transformation, key components of a solid mobile strategy and explored how brands are effectively driving mobile change throughout the enterprise.
Internet of @ThingsExpo, taking place November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 19th Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The Internet of Things (IoT) is the most profound change in personal and enterprise IT since the creation of the Worldwide Web more than 20 years ago. All major researchers estimate there will be tens of billions devices - comp...
DevOps at Cloud Expo, taking place Nov 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 19th Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The widespread success of cloud computing is driving the DevOps revolution in enterprise IT. Now as never before, development teams must communicate and collaborate in a dynamic, 24/7/365 environment. There is no time to wait for long dev...
Almost two-thirds of companies either have or soon will have IoT as the backbone of their business in 2016. However, IoT is far more complex than most firms expected. How can you not get trapped in the pitfalls? In his session at @ThingsExpo, Tony Shan, a renowned visionary and thought leader, will introduce a holistic method of IoTification, which is the process of IoTifying the existing technology and business models to adopt and leverage IoT. He will drill down to the components in this fra...
Data is the fuel that drives the machine learning algorithmic engines and ultimately provides the business value. In his session at Cloud Expo, Ed Featherston, a director and senior enterprise architect at Collaborative Consulting, will discuss the key considerations around quality, volume, timeliness, and pedigree that must be dealt with in order to properly fuel that engine.
There is growing need for data-driven applications and the need for digital platforms to build these apps. In his session at 19th Cloud Expo, Muddu Sudhakar, VP and GM of Security & IoT at Splunk, will cover different PaaS solutions and Big Data platforms that are available to build applications. In addition, AI and machine learning are creating new requirements that developers need in the building of next-gen apps. The next-generation digital platforms have some of the past platform needs a...
19th Cloud Expo, taking place November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA, will feature technical sessions from a rock star conference faculty and the leading industry players in the world. Cloud computing is now being embraced by a majority of enterprises of all sizes. Yesterday's debate about public vs. private has transformed into the reality of hybrid cloud: a recent survey shows that 74% of enterprises have a hybrid cloud strategy. Meanwhile, 94% of enterpri...
SYS-CON Events announced today Telecom Reseller has been named “Media Sponsor” of SYS-CON's 19th International Cloud Expo, which will take place on November 1–3, 2016, at the Santa Clara Convention Center in Santa Clara, CA. Telecom Reseller reports on Unified Communications, UCaaS, BPaaS for enterprise and SMBs. They report extensively on both customer premises based solutions such as IP-PBX as well as cloud based and hosted platforms.
Pulzze Systems was happy to participate in such a premier event and thankful to be receiving the winning investment and global network support from G-Startup Worldwide. It is an exciting time for Pulzze to showcase the effectiveness of innovative technologies and enable them to make the world smarter and better. The reputable contest is held to identify promising startups around the globe that are assured to change the world through their innovative products and disruptive technologies. There w...
With so much going on in this space you could be forgiven for thinking you were always working with yesterday’s technologies. So much change, so quickly. What do you do if you have to build a solution from the ground up that is expected to live in the field for at least 5-10 years? This is the challenge we faced when we looked to refresh our existing 10-year-old custom hardware stack to measure the fullness of trash cans and compactors.
The emerging Internet of Everything creates tremendous new opportunities for customer engagement and business model innovation. However, enterprises must overcome a number of critical challenges to bring these new solutions to market. In his session at @ThingsExpo, Michael Martin, CTO/CIO at nfrastructure, outlined these key challenges and recommended approaches for overcoming them to achieve speed and agility in the design, development and implementation of Internet of Everything solutions wi...
Cloud computing is being adopted in one form or another by 94% of enterprises today. Tens of billions of new devices are being connected to The Internet of Things. And Big Data is driving this bus. An exponential increase is expected in the amount of information being processed, managed, analyzed, and acted upon by enterprise IT. This amazing is not part of some distant future - it is happening today. One report shows a 650% increase in enterprise data by 2020. Other estimates are even higher....
Today we can collect lots and lots of performance data. We build beautiful dashboards and even have fancy query languages to access and transform the data. Still performance data is a secret language only a couple of people understand. The more business becomes digital the more stakeholders are interested in this data including how it relates to business. Some of these people have never used a monitoring tool before. They have a question on their mind like “How is my application doing” but no id...
The 19th International Cloud Expo has announced that its Call for Papers is open. Cloud Expo, to be held November 1-3, 2016, at the Santa Clara Convention Center in Santa Clara, CA, brings together Cloud Computing, Big Data, Internet of Things, DevOps, Digital Transformation, Microservices and WebRTC to one location. With cloud computing driving a higher percentage of enterprise IT budgets every year, it becomes increasingly important to plant your flag in this fast-expanding business opportuni...
Smart Cities are here to stay, but for their promise to be delivered, the data they produce must not be put in new siloes. In his session at @ThingsExpo, Mathias Herberts, Co-founder and CTO of Cityzen Data, will deep dive into best practices that will ensure a successful smart city journey.
SYS-CON Events announced today that 910Telecom will exhibit at the 19th International Cloud Expo, which will take place on November 1–3, 2016, at the Santa Clara Convention Center in Santa Clara, CA. Housed in the classic Denver Gas & Electric Building, 910 15th St., 910Telecom is a carrier-neutral telecom hotel located in the heart of Denver. Adjacent to CenturyLink, AT&T, and Denver Main, 910Telecom offers connectivity to all major carriers, Internet service providers, Internet backbones and ...
Identity is in everything and customers are looking to their providers to ensure the security of their identities, transactions and data. With the increased reliance on cloud-based services, service providers must build security and trust into their offerings, adding value to customers and improving the user experience. Making identity, security and privacy easy for customers provides a unique advantage over the competition.
I wanted to gather all of my Internet of Things (IOT) blogs into a single blog (that I could later use with my University of San Francisco (USF) Big Data “MBA” course). However as I started to pull these blogs together, I realized that my IOT discussion lacked a vision; it lacked an end point towards which an organization could drive their IOT envisioning, proof of value, app dev, data engineering and data science efforts. And I think that the IOT end point is really quite simple…
Personalization has long been the holy grail of marketing. Simply stated, communicate the most relevant offer to the right person and you will increase sales. To achieve this, you must understand the individual. Consequently, digital marketers developed many ways to gather and leverage customer information to deliver targeted experiences. In his session at @ThingsExpo, Lou Casal, Founder and Principal Consultant at Practicala, discussed how the Internet of Things (IoT) has accelerated our abil...
Is the ongoing quest for agility in the data center forcing you to evaluate how to be a part of infrastructure automation efforts? As organizations evolve toward bimodal IT operations, they are embracing new service delivery models and leveraging virtualization to increase infrastructure agility. Therefore, the network must evolve in parallel to become equally agile. Read this essential piece of Gartner research for recommendations on achieving greater agility.