mpackdb
All repositories: gitoria
9.9 KB
# MPackDB Architecture## OverviewMPackDB is a fast, local, append-only JSON database that uses MessagePack serialization. It's designed for Node.js/Bun applications that need a simple, file-based database with optional indexing capabilities and smaller file sizes compared to BSON.## Core Components### 1. MPackDB Class (`src/MPackDB.js`)The main database class that handles all CRUD operations and coordinates between components.**Key Responsibilities:**- Database initialization and lifecycle management- CRUD operations (insert, update, delete, find)- File locking for concurrent access- Metadata persistence- Compaction of deleted records**Key Properties:**- `_dataPath`: Path to the `.mpack` data file- `_dataStream`: Write stream for append-only operations- `_meta`: Metadata object containing `nextId` and `deleted` offsets- `_indexManager`: Optional IndexManager instance for indexed queries- `_primaryKey`: Name of the primary key field- `_primaryKeyType`: Type of primary key (NUMBER, UUID, STRING)### 2. IndexManager Class (`src/IndexManager.js`)Manages binary search indexes for fast lookups on indexed fields.**Key Responsibilities:**- Building and maintaining indexes from data files- Binary search on disk-based index files- Delta indexes (in-memory changes not yet persisted)- Tombstone tracking for deleted records- Auto-persistence of indexes**Index File Format:**```key,offset,lengthkey,offset,length...```Each line represents an index entry where:- `key`: The indexed field value- `offset`: Byte offset in the data file- `length`: Length of the record in bytes**Index Types:**- **NUMERIC**: Numeric comparison for sorting/searching- **LEXICAL**: String comparison for sorting/searching### 3. Cursor Class (`src/Cursor.js`)Provides an async iterable interface for query results.**Key Responsibilities:**- Lazy evaluation of queries- Support for different iteration modes (record, offset, mixed, raw)- Integration with IndexManager for indexed queries- Filtering with query functions### 4. MessagePack Utilities (`src/mpack.js`)Wrapper around `msgpackr` with custom utilities.**Key Exports:**- `serialize()`: Encode JavaScript objects to MessagePack binary- `deserialize()`: Decode MessagePack binary to JavaScript objects- `uuid()`: Generate sortable 12-char base36 unique IDs (9 timestamp + 3 random)- `PrimaryKeyType`: Enum for primary key types- `IndexType`: Enum for index types## Data Flow### Insert Operation```1. User calls db.insert(record)2. Acquire file lock3. Auto-generate primary key if needed (numeric/UUID)4. Serialize record to MessagePack5. Prepend 4-byte size header6. Append to data file via write stream7. Add entry to IndexManager (if indexes enabled)8. Persist metadata (if nextId changed)9. Release file lock10. Return primary key or full record```### Find Operation (Indexed)```1. User calls db.find(primaryKeyValue)2. Cursor created with query function3. IndexManager performs binary search on index file4. Check delta indexes for recent changes5. Check tombstones for deleted records6. Read record from data file at found offset7. Yield record to user```### Find Operation (Non-Indexed)```1. User calls db.find(queryFn)2. Cursor created with query function3. Stream through entire data file4. Read 4-byte size header5. Read MessagePack data based on size6. Deserialize each record7. Apply query function filter8. Skip deleted records (check metadata.deleted)9. Yield matching records to user```### Update Operation```1. User calls db.update(query, newData)2. Find matching records (using find)3. For each match:a. Mark old record as deletedb. Insert new record with same primary key4. Persist metadata with new deleted offsets```### Delete Operation```1. User calls db.delete(query)2. Find matching records3. Add offsets to metadata.deleted array4. Remove from indexes (if enabled)5. Persist metadata```### Compact Operation```1. Acquire file lock2. Create temporary data file3. Stream through all records4. Write only non-deleted records to temp file5. Atomically rename temp file to replace original6. Clear metadata.deleted array7. Rebuild all indexes from new file8. Release file lock```## File Structure```data/├── users.mpack # Main data file (MessagePack records)├── users.meta.json # Metadata (nextId, deleted offsets)├── users.lock # Lock file (contains PID)├── users.id.txt # Index file for 'id' field├── users.email.txt # Index file for 'email' field└── ...```## Serialization FormatMPackDB uses MessagePack format from the `msgpackr` npm package with a custom size header:```[4 bytes: size][MessagePack data]```Each record is prefixed with a 4-byte little-endian integer indicating the size of the MessagePack data (not including the size prefix itself).**Why the size header?**- MessagePack doesn't include record boundaries in the format- The size header allows streaming reads without parsing the entire file- Enables skipping deleted records efficiently- Matches the pattern used in BSON for consistency## Locking MechanismMPackDB uses file-based locking to prevent concurrent writes:1. Before any write operation, create `{dbPath}.lock` file with `wx` flag (exclusive)2. Write current process PID to lock file3. If lock exists, wait 100ms and retry4. After operation completes, delete lock fileThis ensures only one process can write at a time while allowing multiple readers.## Index Persistence StrategyIndexes use a two-tier approach:### Disk Indexes- Sorted index files on disk- Binary searchable for O(log n) lookups- Rebuilt during compaction### Delta Indexes (In-Memory)- Track changes since last persistence- Checked before disk indexes- Auto-persisted based on:- Time interval (default: 60 seconds)- Change threshold (default: 1000 operations)### Tombstones- Track deleted records in memory- Prevent returning deleted records from disk indexes- Cleared during compaction## Performance Characteristics### Time Complexity- **Insert**: O(1) for append, O(log n) for index update- **Find by primary key (indexed)**: O(log n) binary search- **Find with query function**: O(n) full scan- **Update**: O(log n) find + O(1) insert- **Delete**: O(log n) find + O(1) mark- **Compact**: O(n) full scan + O(n log n) index rebuild### Space Complexity- Data file grows with inserts (append-only)- Deleted records remain until compaction- Index files: O(n) per indexed field- Delta indexes: O(m) where m = changes since last persist### File Size ComparisonMessagePack typically produces **15-20% smaller files** than BSON for the same data:- More compact integer encoding- Smaller string overhead- Efficient array/map encoding## Concurrency Model- **Single-writer, multiple-reader** via file locking- Writes are serialized through lock file- Reads can happen concurrently (no locks needed)- Index persistence happens asynchronously but safely## Primary Key Types### NUMBER (PrimaryKeyType.NUMBER)- Auto-incremented integer- Stored in metadata.nextId- Prefix syntax: `*id`### UUID (PrimaryKeyType.UUID)- Sortable base36 unique ID (9-char timestamp + 3-char random)- 12-character string format- Prefix syntax: `@id`### STRING (PrimaryKeyType.STRING)- User-provided string- No auto-generation- Default (no prefix)## Design Decisions### Why Append-Only?- **Fast writes**: No seeking, just append- **Crash safety**: Partial writes don't corrupt existing data- **Simple implementation**: No complex update-in-place logic### Why MessagePack?- **Smaller files**: 15-20% smaller than BSON on average- **Fast serialization**: Comparable or faster than BSON- **Wide language support**: Available in many programming languages- **Simple format**: Easy to implement and debug- **No external binary dependencies**: Pure JavaScript implementation### Why Custom Size Header?- MessagePack doesn't define record boundaries- Enables efficient streaming without full deserialization- Allows skipping deleted records quickly- Consistent with BSON's approach### Why File-Based Locking?- **Simple**: No external dependencies- **Cross-process**: Works across multiple Node.js processes- **Portable**: Works on all platforms### Why Binary Search Indexes?- **Disk-friendly**: Can search large indexes without loading into memory- **Simple format**: Plain text, easy to debug- **Fast lookups**: O(log n) for indexed queries## MessagePack vs BSON### Advantages of MessagePack- **Smaller files**: 15-20% size reduction- **Faster reads**: Simpler format, less parsing overhead- **Pure JavaScript**: No native dependencies- **Smaller library**: ~15KB vs ~173KB for BSON### Advantages of BSON- **ObjectId type**: Built-in unique identifier type- **Date precision**: Millisecond timestamps- **Binary data**: Native binary type- **MongoDB compatibility**: Direct compatibility with MongoDB### When to Choose MPackDB- File size is a concern- Pure JavaScript dependencies preferred- Don't need MongoDB compatibility- Want faster read performance### When to Choose BsonDB- Need ObjectId primary keys- MongoDB compatibility desired- Working with binary data- Need precise date/time handling## Limitations1. **Single-writer**: Only one write operation at a time2. **No transactions**: Operations are not atomic across multiple records3. **No query language**: Must use JavaScript functions for complex queries4. **Compaction required**: Deleted records consume space until compaction5. **Index overhead**: Each index doubles storage for that field6. **No schema validation**: Records can have any structure7. **No ObjectId type**: Must use UUIDs or numeric IDs## Future Improvements- Batch insert operations- Async compaction (background process)- Query optimizer for complex filters- Compression support (MessagePack supports extensions)- Replication/backup utilities- Schema validation layer- Custom MessagePack extension types
Branches
- mastermain branch