International Edition
Latest News
Technology

Reverse-Engineering the Intel 8087 FPU: How FPTAN Works

Reverse-engineering efforts on the historic Intel 8087 floating-point unit reveal how chip designers balanced trigonometric accuracy with hardware constraints through a hybrid math approach. Software engineer Ken Shirriff and collaborators detailed the microcode and algorithmic structure behind the…

Reverse-Engineering the Intel 8087 FPU: How FPTAN Works

Reverse-engineering efforts on the historic Intel 8087 floating-point unit reveal how chip designers balanced trigonometric accuracy with hardware constraints through a hybrid math approach. Software engineer Ken Shirriff and collaborators detailed the microcode and algorithmic structure behind the processor’s tangent instruction, showing how engineers combined the CORDIC method with polynomial approximation to achieve 64-bit precision.

Algorithm Selection in Early Math Coprocessors

Processors lacking dedicated hardware floating-point support—such as the 6502 or Z80—typically rely on algorithms like CORDIC to compute trigonometric functions. This approach requires only basic hardware features including addition, subtraction, bit-shifting, and look-up tables. However, scaling CORDIC to high bit-widths demands significantly larger tables and longer execution times. For the 8087 coprocessor, Intel engineers designed a hybrid architecture to maximize both speed and accuracy without expanding look-up table requirements.

Reverse-Engineering the Intel 8087 FPU: How FPTAN Works

The Hybrid Path to 64-Bit Precision

The analysis of the FPTAN instruction demonstrates a two-stage calculation process. The system first computes the initial 16 bits using the CORDIC method. It then switches to a Padé approximant technique, which uses the ratio of two polynomials to resolve the remaining bits. Because CORDIC reduces the remainder to a very small value before the handoff, the subsequent polynomial approximation runs quickly while maintaining high accuracy.

Timing breakdowns for the operation show that the instruction spends roughly 33 percent of its cycles on CORDIC pseudo-division, 47 percent on CORDIC pseudo-multiplication, and just 15 percent on the polynomial approximation, with a five percent overhead. This division of labor allowed the 8087 to deliver high-precision results while sidestepping the performance penalties of a pure CORDIC pipeline.

Architectural Shifts in Later Processor Generations

As processor design progressed, hardware architectures moved away from the CORDIC approach. Intel abandoned CORDIC entirely with the introduction of the Pentium series, as scaling the method to larger bit-widths proved too slow for advancing clock speeds. Subsequent integration of SIMD instruction sets further reduced the role of the x87 instruction set architecture. Despite these shifts, the reverse-engineered microcode of the 8087 illustrates the foundational design choices that established hardware floating-point performance.

Fake $9 Chinese Intel 8087 chip from eBay. Will it work?
About the author: Anika Shah - Technology

MSc in Computer Science, senior reporter. Anika focuses on AI ethics, cybersecurity, and emerging hardware—frequently moderating panels at CES and Web Summit. “Anika Shah decodes tech breakthroughs and startup disruption shaping tomorrow’s digital landscape.”