
Understanding Automatic Storage Tiering
Fast Storage and Large Storage Solve Different Problems
Solid-state drives and traditional hard drives provide very different advantages. SSDs can respond to storage requests extremely quickly because they do not rely on mechanical heads moving across rotating platters. Hard drives, however, can provide large amounts of capacity economically.
A system built entirely from SSDs can provide excellent performance but at considerably greater cost per unit of storage. A system relying entirely on hard drives can provide substantial capacity while giving frequently accessed workloads much slower response times.
Storage tiering offers another approach. Instead of requiring every piece of information to reside on the same type of device, the storage system can combine fast and high-capacity media and place data according to how actively it is being used.
Not Every Byte Needs the Fastest Drive
Much of the information stored on a large system may be accessed infrequently. Reserving expensive high-performance storage for the portions being used heavily can provide much of the performance benefit without requiring the entire dataset to reside on SSDs.
Storage Activity Is Not Distributed Evenly
A large volume can contain enormous amounts of information while only a relatively small portion receives frequent requests at any particular time.
Operating-system files, active databases, frequently used virtual-machine data, application components, and current projects may generate substantial storage activity. Older archives and rarely accessed information can remain almost untouched for long periods.
Storage systems commonly describe these patterns using the terms hot and cold. Hot data is accessed frequently. Cold data receives comparatively little activity.
Hot Data
Frequently accessed portions of stored information benefit strongly from low-latency, high-performance storage.
Cold Data
Information that is rarely accessed can remain on larger and less expensive storage without consuming limited high-speed capacity.
The Classification Can Change Over Time
Hot and cold are not permanent properties of a file. They describe activity.
A project that has been untouched for months can suddenly become active again. An operating-system component that was accessed constantly during an installation may receive little attention afterward. Database regions can become busy or quiet as workloads change.
A useful tiering system therefore needs to observe storage activity continuously rather than permanently assigning information based only on what happened when the file was first created.
Temperature Describes Behavior
Calling data hot or cold does not describe its importance or file type. It describes how frequently that portion of storage is being accessed during the period being measured.
SSDs and Hard Drives Can Participate in the Same Virtual Disk
Storage Spaces can combine physical disks into storage pools and create virtual disks from the available capacity. Storage tiering extends that model by recognizing different media types inside the pool.
A tiered virtual disk can contain a high-performance SSD tier and a larger standard hard-drive tier. To the workload using the resulting volume, those layers can appear as part of the same logical storage resource.
The operating system handles the placement of information between the underlying media rather than requiring applications to maintain separate SSD and HDD locations.
One Volume Can Span Different Storage Technologies
The logical volume presented to an application does not have to correspond to one physical disk or even one type of storage media. Virtualization allows SSDs and hard drives to contribute different characteristics to the same storage system.
The SSD Tier Does Not Need to Hold the Entire Dataset
The performance advantage of tiering comes from concentrating limited SSD capacity where it can have the greatest effect.
Suppose several terabytes of information are stored on a volume while only a fraction of that data is accessed intensively. Moving the active portions to SSDs can accelerate many storage requests even though most of the total capacity remains on hard drives.
This changes the economics of high-performance storage. The system does not need enough flash capacity to contain every file before SSD performance can benefit the workload.
The objective is not to make every byte reside on the fastest device. It is to make the fastest device available to the bytes that currently benefit from it most.
The System Measures Which Data Is Being Accessed
Automatic tiering requires evidence about storage activity. The system observes which portions of data receive frequent access and uses that information when deciding what belongs on the faster tier.
This differs from manually deciding that one entire folder should live on an SSD while another should live on a hard drive. The placement decision can occur below the level of complete files.
That finer granularity is important because a very large file may contain some regions that are used constantly and others that are rarely touched.
The Entire File Does Not Have to Move
Storage Spaces can perform tiering at a sub-file level. Frequently accessed portions can move toward faster storage without requiring every part of a large file to consume SSD capacity.
Frequently Accessed Regions Can Move Up to the SSD Tier
When activity measurements show that a portion of stored information is being accessed frequently, that region can be moved from the standard hard-drive tier to the faster SSD tier.
Once there, subsequent requests can benefit from the lower latency and high random-access performance associated with solid-state storage.
The application does not need to know that the physical location changed. It continues accessing the same logical data while the storage layer manages where that information resides underneath.
The Optimization Happens Below the Application
Software using the volume does not ordinarily need separate paths for fast and slow storage. The tiering mechanism manages physical placement while preserving the logical storage presented to the workload.
Cold Data Can Return to the Hard-Drive Tier
SSD capacity is valuable precisely because it is limited. If information moved to the fast tier remained there forever, the SSD portion would eventually fill with data that was once active but no longer receives frequent requests.
Automatic tiering therefore works in both directions. Information that becomes less active can move back toward the standard tier, releasing SSD capacity for data that has become hotter.
The storage layout can consequently adapt as workload behavior changes.
Does Frequently Used Data Stay on the SSD Forever?
Not necessarily. Placement is based on activity. Information that becomes less frequently accessed can move toward the hard-drive tier while newly active portions take advantage of the faster storage.
The Working Set Can Be Much Smaller Than the Stored Dataset
The concept of a working set helps explain why tiering can be effective. A system may store several terabytes while actively working with only a relatively small portion during normal operation.
If that active portion fits within the SSD tier, many important storage requests can receive flash-level performance even though the majority of information remains on mechanical drives.
This is especially useful for workloads containing large datasets with concentrated areas of frequent activity.
Capacity and Activity Are Different
The amount of information stored does not tell how much of it needs high performance at the same moment. Tiering takes advantage of the difference between total dataset size and the smaller collection of data receiving frequent access.
SSDs Can Also Buffer Random Write Activity
Reading frequently accessed data is not the only workload that benefits from flash storage. Small random writes can create substantial performance pressure on mechanical disks because the drive must repeatedly reposition its heads to different physical locations.
A write-back cache can use SSD capacity to absorb small random writes quickly and later transfer the information to hard-drive storage.
This allows the workload to complete certain write operations with lower latency while the storage system handles slower placement onto mechanical media afterward.
Caching and Tiering Are Related but Different
Storage tiering decides where data should reside according to activity, while a write-back cache temporarily absorbs writes on faster storage before they are committed to their longer-term location. Windows Server 2012 R2 introduced both capabilities for Storage Spaces.
Mechanical Drives Struggle With Random Access
A hard drive can transfer sequential information efficiently once its heads are positioned correctly. Random workloads are more difficult because each request may require mechanical movement to another location on the platter.
SSDs do not have that mechanical seek process. Accessing information from widely separated logical locations therefore does not impose the same physical movement penalty.
This difference is one reason a relatively small amount of solid-state storage can have an outsized effect on workloads dominated by random access.
Sequential Activity
Large contiguous transfers can make good use of hard-drive throughput once the mechanical device is positioned for the operation.
Random Activity
Requests scattered throughout storage can force mechanical drives to perform repeated seek operations that increase latency.
Solid-State Access
Flash storage eliminates mechanical head movement, making it particularly effective for workloads containing frequent random requests.
Specific Files Can Be Pinned to a Storage Tier
Automatic placement is useful when workload activity should determine where information resides, but some situations require predictable placement.
Storage tiering can allow administrators to pin particular files to a chosen tier. A file that should remain on high-performance storage can therefore be assigned to the SSD tier rather than depending entirely on automatic activity analysis.
The reverse can also be useful when a large file should remain on standard storage even if temporary activity might otherwise encourage promotion.
Automatic Optimization Can Still Allow Policy
Tiering can respond dynamically to measured activity while administrators retain the ability to keep selected files on a particular tier when workload requirements justify fixed placement.
The Fast Tier Cannot Compensate for Every Storage Bottleneck
Adding SSD capacity to a tiered pool does not guarantee that every workload becomes fast.
The active dataset may exceed the available SSD tier. A workload may perform large sequential operations that gain less from flash acceleration. Processor, memory, network, controller, or application limitations can also become the dominant bottleneck.
Storage performance therefore has to be evaluated as part of the complete system rather than attributed to one device type.
An SSD Tier Is Not Unlimited Performance
The benefit depends on workload behavior, the size of the active dataset, available flash capacity, storage layout, and the performance of the remaining system. Tiering improves where data resides; it cannot remove unrelated bottlenecks.
Virtualization Does Not Remove Hardware Failure
Tiering changes the placement of information, but the underlying data still resides on physical devices that can fail.
Storage Spaces can combine tiering with resilient layouts so that performance and fault tolerance are considered together. The required number and arrangement of drives depend on the selected storage design.
An SSD failing in the fast tier or a hard drive failing in the capacity tier therefore remains a hardware event that the storage system must handle according to its resiliency configuration.
Performance and Resilience Are Separate Decisions
Storage tiers determine which class of media should hold particular data. Mirroring, parity, and other resiliency mechanisms determine how the storage system responds when physical devices fail.
The Best Storage Location Became Dynamic
Traditional storage planning often required administrators to decide in advance which data belonged on expensive fast disks and which belonged on larger slower disks. Those decisions could become outdated as workloads changed.
Automatic storage tiering moved part of that decision into the storage system itself. Activity could be measured, frequently accessed portions promoted to SSDs, and colder portions returned to hard drives as their usefulness to the fast tier declined.
This allowed the physical placement of data to evolve without requiring applications or users to reorganize their files continually.
Fast Storage Could Be Concentrated Where It Mattered Most
The larger idea behind storage tiering is resource allocation. SSD capacity is valuable because it provides excellent performance, but that performance does not have to be distributed equally across every stored byte.
By identifying the portions receiving the greatest activity, a tiered system can concentrate fast storage around the current working set while retaining economical hard-drive capacity for the much larger collection of colder information.
As the workload changes, the placement can change with it. Data that becomes active can move upward. Data that becomes quiet can move downward. The logical volume remains the same while its physical organization evolves underneath.
Storage Performance Became Less Dependent on Permanent Placement
Automatic tiering represented a shift from static storage design toward adaptive storage behavior. Instead of requiring an administrator to predict permanently which information deserved the fastest devices, the system could observe actual use and make placement decisions from that evidence.
The combination of SSD responsiveness and hard-drive capacity made it possible to build storage systems that balanced performance and cost without treating those goals as mutually exclusive.
The important change was not simply putting an SSD beside a hard drive. It was giving the storage layer enough intelligence to decide which portions of the workload should benefit from each one.