We have a collection of log data, where each document in the collection is identified by a MAC address and a calendar day. Basically:

{
  _id: <generated>,
  mac: <string>,
  day: <date>,
  data: [ "value1", "value2" ]
}

Every five minutes, we append a new log entry to the data array within the current day's document. The document rolls over at midnight UTC when we create a new document for each MAC.

We've noticed that IO, as measured by bytes written, increases all day long, and then drops back down at midnight UTC. This shouldn't happen because the rate of log messages is constant. We believe that the unexpected behavior is due to Mongo moving documents, as opposed to updating their log arrays in place. For what it's worth, stats() shows that the paddingFactor is 1.0299999997858227.

Several questions:

  1. Is there a way to confirm whether Mongo is updating in place or moving? We see some moves in the slow query log, but this seems like anecdotal evidence. I know I can db.setProfilingLevel(2), then db.system.profile.find(), and finally look for "moved:true", but I'm not sure whether it's ok to do this on a busy production system.
  2. The size of each document is very predictable and regular. Assuming that mongo is doing a lot of moves, what's the best way to figure out why isn't Mongo able to presize more accurately? Or to make Mongo presize more accurately? Assuming that the above description of the problem is right, tweaking the padding factor does not seem like it would do the trick.
  3. It should be easy enough for me to presize the document and remove any guesswork from Mongo. (I know the padding factor docs say that I shouldn't have to do this, but I just need to put this issue behind me.) What's the best way to presize a document? It seems simple to writ
Edit
Report