KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
I have an application where I am reading and writing small blocks of data (a few hundred bytes) hundreds of millions of times. I'd like to generate a compression dictionary based on an example data file and use that dictionary forever as I read and write the small blocks. I'm leaning toward the LZW compression algorithm. The Wikipedia page ( http://en.wikipedia.org/wiki/Lempel-Ziv-Welch ) lists pseudocode for compression and decompression. It looks fairly straightforward to modify it such that the dictionary creation is a separate block of code. So I have two questions: Am I on the right track or is there a better way? Why does the LZW algorithm add to the dictionary during the decompression step? Can I omit that, or would I lose efficiency in my dictionary? Thanks. Update: Now I'm thinking the ideal case be to find a library that lets me store the dictionary separate from the compressed data. Does anything like that exist? Update: I ended up taking the code at http://www.enusbaum.com/blog/2009/05/22/example-huffman-compression-routine-in-c and adapting it. I am Chris in the comments on that page. I emailed my mods back to that blog author, but I haven't heard back yet. The compression rates I'm seeing with that code are not at all impressive. Maybe that is due to the 8-bit tree size. Update: I converted it to 16 bits and the compression is better. It's also much faster than the original code. using System; using System.Collections.Generic; using System.Linq; using System.Text; using System.IO; namespace Book.Core { public class Huffman16 { private readonly double log2 = Math.Log(2); private List<Node> HuffmanTree =
Tags (comma-separated)
Save Edits
Cancel