KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
When debugging, I frequently stepped into the handwritten assembly implementation of memcpy and memset. These are usually implemented using streaming instructions if available, loop unrolled, alignment optimized, etc... I also recently encountered this 'bug' due to memcpy optimization in glibc . The question is: why can't the hardware manufacturers (Intel, AMD) optimize the specific case of rep stos and rep movs to be recognized as such, and do the fastest fill and copy as possible on their own architecture?
Tags (comma-separated)
Save Edits
Cancel