Cost.
(Note that memcpy is coming to ARM, see below.)
The cost of optimizing memcpy in your C library is fairly minimal, maybe a few weeks of developer time here and there. You'll have to make a new version every several years or so when processor features change enough to warrant a rewrite. For example, GNU's glibc and Apple's libSystem both have a memcpy which is specifically optimized for SSE3.
The cost of optimizing in hardware is much higher. Not only is it more expensive in terms of developer costs (designing a CPU is vastly more difficult than writing user-space assembly code), but it would increase the transistor count of the processor. That could have a number of negative effects:
- Increased power consumption
- Increased unit cost
- Increased latency for certain CPU subsystems
- Lower maximum clock speed
In theory, it could have an overall negative impact on both performance and unit cost.
Maxim: Don't do it in hardware if the software solution is good enough.
But, memcpy is coming to ARM. As processors get larger and larger, the incremental cost of adding more instructions gets lower and lower, relative to the existing cost of a core. From Arm A-Profile Architecture Developments 2021:
To address these concerns the 2021 extensions introduce new instructions specifically targeting the memcpy() and memset() family of functions.
The key concerns mentioned are that the complicated software implementations of memcpy, while fast, may need to be rewritten for different microarchitectures in order to get better performance. They also need to account for alignment and size in different ways. Having a fa
answered 2012-01-14T00:28:33.910