Five-relay circuit board with damaged second relay marked with X
A circuit board equipped with five relays positioned next to each other has a failed second relay, identified by the X marking. The damaged relay caused the board to malfunction and required replacement. This repair image is an independent work sample and is not an illustration of the educational subject discussed below.

Graphics processors had spent years becoming larger, faster, and more complex, but improving performance did not always mean simply adding more processing hardware or allowing a chip to consume more electricity. The manufacturing technology underneath the processor could change what was practical.

AMD’s Polaris generation demonstrated this particularly well. Its Radeon graphics processors moved away from the 28-nanometer manufacturing process used by earlier generations and adopted a 14-nanometer FinFET process.

The change affected much more than the physical dimensions associated with manufacturing. Smaller FinFET transistors could operate with reduced leakage and lower voltage, giving GPU designers additional room to balance clock speed, performance, heat, and power consumption.

The Transistor Changed Shape

Traditional planar transistors place the controlling gate across a relatively flat channel. As semiconductor features became smaller, controlling electrical leakage through that arrangement became increasingly difficult.

FinFET manufacturing changed the geometry. Instead of relying on a flat channel, the transistor channel rises into a narrow fin and the gate surrounds more of it. That gives the gate greater electrical control over whether current should flow.

Process
14nm FinFET
Architecture
Fourth-generation GCN
Polaris 10 transistors
5.7 billion
Polaris 10 die area
About 230 mm²

For a graphics processor containing billions of transistors, improvements at the level of each transistor could accumulate into a significant change across the complete chip.

Efficiency Created More Design Choices

A more efficient transistor does not automatically make every graphics card faster. Instead, it changes the engineering limits within which the processor can be designed.

If a transistor can operate at a required frequency while consuming less power, the processor may perform the same workload more efficiently. Alternatively, some of that electrical and thermal headroom can be used to increase operating frequency or support additional processing activity.

A smaller manufacturing process mattered because it changed how much useful graphics work could fit inside a practical power and heat budget.

This was especially important for graphics processors because a high-performance GPU can place a substantial continuous load on both the power-delivery circuitry and the cooling system.

Polaris Was Still Built on Familiar Foundations

The move to FinFET did not mean AMD discarded its previous graphics architecture and started over. Polaris remained part of the Graphics Core Next family.

Its Compute Units retained the broad organization familiar from earlier GCN processors, while AMD made targeted changes around them. Improvements included instruction prefetching, cache behavior, geometry processing, memory compression, and scheduling.

That evolutionary approach allowed the manufacturing transition and architectural refinements to work together. Some efficiency came from the semiconductor process itself, while other improvements came from changing how the GPU handled its workload.

Discarding Unnecessary Geometry Saved Work

A graphics processor receives geometric information describing the objects that may eventually appear in a rendered frame. Not every primitive entering the pipeline will ultimately contribute a visible pixel.

Polaris included improvements to geometry processing that could identify and discard certain primitives before they traveled farther through the rasterization process.

Removing work that cannot affect the final image is valuable because every unnecessary operation consumes processing time, memory bandwidth, and electrical power. Efficiency therefore depended not only on performing operations faster, but also on avoiding operations that did not need to be performed at all.

Efficiency could come from doing less unnecessary work

A processor does not have to accelerate every operation to improve throughput. Detecting work that can be safely discarded before later stages of the graphics pipeline can free resources for operations that actually contribute to the rendered frame.

Memory Traffic Was Another Place to Save Energy

Graphics processors continuously move large amounts of information between processing units, caches, and external graphics memory. That traffic has a cost. Moving data consumes bandwidth and electrical power.

Polaris improved delta color compression, allowing suitable framebuffer information to be represented more efficiently as it moved through parts of the memory system.

Compression did not mean that the final image had to be visibly degraded. The goal was to reduce the amount of memory traffic required to represent information that could be reconstructed without changing the intended pixel values.

Using available memory bandwidth more efficiently could help the GPU spend less time moving redundant information and more effectively use the physical memory interface already available to it.

The GPU Was Managing Its Work More Carefully

Modern graphics processors do much more than execute one continuous stream of traditional rendering instructions. Different graphics and compute workloads may need to share processing resources, and the GPU has to decide how that work should be scheduled.

Polaris continued AMD’s use of asynchronous compute hardware while refining the mechanisms used to schedule work. Hardware schedulers could assist in distributing tasks and managing compute activity without requiring every scheduling decision to be handled in exactly the same way by software.

Processing

Compute Units performed the shader and compute operations that formed the core of the graphics workload.

Scheduling

Hardware scheduling mechanisms helped coordinate work competing for GPU processing resources.

Data movement

Cache and compression improvements helped reduce unnecessary pressure on external memory bandwidth.

These areas were interconnected. Faster execution could be wasted if the processor stalled waiting for data, while plentiful memory bandwidth could be wasted if scheduling left processing resources idle.

Power Management Could React to the Individual Chip

Semiconductor manufacturing produces small electrical differences between individual chips. Two processors of the same model are designed to meet the same specifications, but their exact voltage and leakage characteristics are not perfectly identical.

Polaris introduced more refined power-management behavior that could account for characteristics of the individual processor rather than relying entirely on broad assumptions intended to cover manufacturing variation.

Better knowledge of the chip’s electrical behavior made it possible to reduce some of the excess margin that would otherwise consume power without contributing useful graphics performance.

Power efficiency was a system of improvements

The manufacturing process provided an important foundation, but the final result also depended on architecture, memory behavior, workload scheduling, voltage control, clock management, and the design of the graphics card surrounding the GPU.

A Smaller Chip Could Compete With Much Larger Designs

Polaris 10 contained approximately 5.7 billion transistors in a die of roughly 230 square millimeters. Earlier high-performance GPUs manufactured on 28nm processes could occupy substantially more silicon for a comparable transistor count.

Older 28nm approach

Larger planar transistors required more die area and presented greater challenges with leakage and power consumption as designs became increasingly complex.

14nm FinFET approach

Smaller FinFET transistors allowed billions of devices to fit into less silicon while improving the electrical characteristics available to the GPU designer.

Die size alone did not determine graphics performance, and processors with different architectures cannot be judged simply by comparing their physical area. The contrast nevertheless illustrated how dramatically semiconductor manufacturing could change the density of a modern GPU.

Performance Was No Longer the Only Number That Mattered

The transition represented by Polaris showed why graphics development could not be measured only by the maximum frame rate of the newest card. Performance had to exist within limits imposed by electricity, heat, cooling, physical size, and cost.

A more efficient architecture could make useful graphics performance practical in systems where a hotter and more power-hungry processor would be difficult to accommodate. That mattered not only for desktop graphics cards, but also for compact computers and mobile designs where cooling and battery capacity imposed much tighter restrictions.

FinFET manufacturing did not eliminate those limits. It changed the amount of graphics work engineers could perform before reaching them. For GPU design, that additional efficiency could be just as important as adding more raw processing hardware.