KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
For one of my projects, I have to enter a big-ish collection of events into a database for later processing and I am trying to decide which DBMS would be best for my purpose. I have: About 400,000,000 discrete events at the moment About 600 GB of data that will be stored in the DB These events come in a variety of formats, but I estimate the count of individual attributes to be about 5,000. Most events only contain values for about 100 attributes each. The attribute values are to be treated as arbitrary strings and, in some cases, integers. The events will eventually be consolidated into a single time series. While they do have some internal structure, there are no references to other events, which - I believe - means that I don't need an object DB or some ORM system. My requirements: Open source license - I may have to tweak it a bit. Scalability by being able to expand to multiple servers, although only one system will be used at first. Fast queries - updates are not that critical. Mature drivers/bindings for C/C++, Java and Python. Preferrably with a license that plays well with others - I'd rather not commit myself to anything because of a technical decision. I think that most DB drivers do not have a problem here, but it should be mentioned, anyway. Availability for Linux. It would be nice, but not necessary, if it was also available for Windows My ideal DB for this would allow me to retrieve all the events from a specified time period with a single query. What I have found/considered so far: Postgresql with an increased page size can apparently have up to 6,000 columns in each table. If my estimate of the attribute count is not off, it might do. <a href="http://www.mysql.
Tags (comma-separated)
Save Edits
Cancel