gitoriaLog in with ident

mpackdb

All repositories: gitoria

ReadmeCodePull requestsReleasesTicketsSettings
Commitb8ffc1a0b8ffc1a0release 1.0.5caramboleyob8ffc1a0/ARCHITECTURE.md

11.4 KB

  1. # MPackDB Architecture
  2. ## Overview
  3. MPackDB is a fast, local, append-only JSON database that uses MessagePack serialization. It's designed for Node.js/Bun applications that need a simple, file-based database with optional indexing capabilities and smaller file sizes compared to BSON.
  4. ## Core Components
  5. ### 1. MPackDB Class (`src/MPackDB.js`)
  6. The main database class that handles all CRUD operations and coordinates between components.
  7. **Key Responsibilities:**
  8. - Database initialization and lifecycle management
  9. - CRUD operations (insert, update, delete, find)
  10. - File locking for concurrent access
  11. - Metadata persistence
  12. - Compaction of deleted records
  13. **Key Properties:**
  14. - `_dataPath`: Path to the `.mpack` data file
  15. - `_dataStream`: Write stream for append-only operations
  16. - `_meta`: Metadata object containing `nextId` and `deleted` offsets
  17. - `_indexManager`: Optional IndexManager instance for indexed queries
  18. - `_primaryKey`: Name of the primary key field
  19. - `_primaryKeyType`: Type of primary key (NUMBER, UUID, STRING)
  20. ### 2. IndexManager Class (`src/IndexManager.js`)
  21. Manages binary search indexes for fast lookups on indexed fields.
  22. **Key Responsibilities:**
  23. - Building and maintaining indexes from data files
  24. - Binary search on disk-based index files
  25. - Delta indexes (in-memory changes not yet persisted)
  26. - Tombstone tracking for deleted records
  27. - Auto-persistence of indexes
  28. **Index File Format:**
  29. ```
  30. key,offset,length
  31. key,offset,length
  32. ...
  33. ```
  34. Each line represents an index entry where:
  35. - `key`: The indexed field value
  36. - `offset`: Byte offset in the data file
  37. - `length`: Length of the record in bytes
  38. **Index Types:**
  39. - **NUMERIC**: Numeric comparison for sorting/searching
  40. - **LEXICAL**: String comparison for sorting/searching
  41. - **UNIQUE**: Any index can be marked unique (`!` prefix) to reject duplicates on insert
  42. ### 3. Cursor Class (`src/Cursor.js`)
  43. Provides an async iterable interface for query results.
  44. **Key Responsibilities:**
  45. - Lazy evaluation of queries
  46. - Support for different iteration modes (record, offset, mixed, raw)
  47. - Integration with IndexManager for indexed queries
  48. - Filtering with query functions
  49. ### 4. MessagePack Utilities (`src/mpack.js`)
  50. Wrapper around `msgpackr` with custom utilities.
  51. **Key Exports:**
  52. - `serialize()`: Encode JavaScript objects to MessagePack binary
  53. - `deserialize()`: Decode MessagePack binary to JavaScript objects
  54. - `uuid()`: Generate sortable 12-char base36 unique IDs (9 timestamp + 3 random)
  55. - `PrimaryKeyType`: Enum for primary key types
  56. - `IndexType`: Enum for index types
  57. ## Data Flow
  58. ### Insert Operation
  59. ```
  60. 1. User calls db.insert(record)
  61. 2. Acquire file lock
  62. 3. Auto-generate primary key if needed (numeric/UUID)
  63. 4. Check unique index constraints (skip auto-generated PKs)
  64. 5. Serialize record to MessagePack
  65. 6. Prepend 4-byte size header
  66. 7. Append to data file via write stream
  67. 8. Add entry to IndexManager (if indexes enabled)
  68. 9. Persist metadata (if nextId changed)
  69. 10. Release file lock
  70. 11. Return primary key or full record
  71. ```
  72. ### Find Operation (Indexed)
  73. ```
  74. 1. User calls db.find(primaryKeyValue)
  75. 2. Cursor created with query function
  76. 3. IndexManager performs binary search on index file
  77. 4. Check delta indexes for recent changes
  78. 5. Check tombstones for deleted records
  79. 6. Read record from data file at found offset
  80. 7. Yield record to user
  81. ```
  82. ### Find Operation (Non-Indexed)
  83. ```
  84. 1. User calls db.find(queryFn)
  85. 2. Cursor created with query function
  86. 3. Stream through entire data file
  87. 4. Read 4-byte size header
  88. 5. Read MessagePack data based on size
  89. 6. Deserialize each record
  90. 7. Apply query function filter
  91. 8. Skip deleted records (check metadata.deleted)
  92. 9. Yield matching records to user
  93. ```
  94. ### findByIndex Operation (Sorted Index Walk)
  95. ```
  96. 1. User calls db.findByIndex(field, { from, limit, filter })
  97. 2. Cursor created with _indexWalk option
  98. 3. Binary search on disk to find start position (from key)
  99. 4. Stream index entries block by block from disk
  100. 5. Merge with in-memory delta entries in sorted order
  101. 6. For each entry: read record from data file at offset
  102. 7. Skip deleted records, apply filter
  103. 8. Yield matching records, stop at limit
  104. ```
  105. ### Update Operation
  106. ```
  107. 1. User calls db.update(query, newData)
  108. 2. Find matching records (using find)
  109. 3. For each match:
  110. a. Mark old record as deleted
  111. b. Insert new record with same primary key
  112. 4. Persist metadata with new deleted offsets
  113. ```
  114. ### Delete Operation
  115. ```
  116. 1. User calls db.delete(query)
  117. 2. Find matching records
  118. 3. Add offsets to metadata.deleted array
  119. 4. Remove from indexes (if enabled)
  120. 5. Persist metadata
  121. ```
  122. ### Compact Operation
  123. ```
  124. 1. Acquire file lock
  125. 2. Create temporary data file
  126. 3. Stream through all records
  127. 4. Write only non-deleted records to temp file
  128. 5. Atomically rename temp file to replace original
  129. 6. Clear metadata.deleted array
  130. 7. Rebuild all indexes from new file
  131. 8. Release file lock
  132. ```
  133. ## File Structure
  134. ```
  135. data/
  136. ├── users.mpack # Main data file (MessagePack records)
  137. ├── users.meta.json # Metadata (nextId, deleted offsets)
  138. ├── users.lock # Lock file (contains PID)
  139. ├── users.id.txt # Index file for 'id' field
  140. ├── users.email.txt # Index file for 'email' field
  141. └── ...
  142. ```
  143. ## Serialization Format
  144. MPackDB uses MessagePack format from the `msgpackr` npm package with a custom size header:
  145. ```
  146. [4 bytes: size][MessagePack data]
  147. ```
  148. Each record is prefixed with a 4-byte little-endian integer indicating the size of the MessagePack data (not including the size prefix itself).
  149. **Why the size header?**
  150. - MessagePack doesn't include record boundaries in the format
  151. - The size header allows streaming reads without parsing the entire file
  152. - Enables skipping deleted records efficiently
  153. - Matches the pattern used in BSON for consistency
  154. ## Locking Mechanism
  155. MPackDB uses file-based locking to prevent concurrent writes:
  156. 1. Before any write operation, create `{dbPath}.lock` file with `wx` flag (exclusive)
  157. 2. Write current process PID to lock file
  158. 3. If lock exists, wait 100ms and retry
  159. 4. After operation completes, delete lock file
  160. The lock is **re-entrant**: if a lock is already held (e.g. inside `withLock()`), nested operations (`insert`, `delete`, etc.) increment a depth counter instead of acquiring a new file lock. The file lock is only released when the outermost holder finishes.
  161. ### withLock for Compound Operations
  162. `db.withLock(callback)` acquires the lock for the duration of the callback. All DB operations inside the callback reuse the same lock. This makes compound operations like find-then-insert atomic:
  163. ```javascript
  164. await db.withLock(async () => {
  165. const [user] = await db.find(u => u.email === email);
  166. if (!user) await db.insert({ email });
  167. });
  168. ```
  169. This ensures only one process can write at a time while allowing multiple readers.
  170. ## Index Persistence Strategy
  171. Indexes use a two-tier approach:
  172. ### Disk Indexes
  173. - Sorted index files on disk
  174. - Binary searchable for O(log n) lookups
  175. - Rebuilt during compaction
  176. ### Delta Indexes (In-Memory)
  177. - Track changes since last persistence
  178. - Checked before disk indexes
  179. - Auto-persisted based on:
  180. - Time interval (default: 60 seconds)
  181. - Change threshold (default: 1000 operations)
  182. ### Tombstones
  183. - Track deleted records in memory
  184. - Prevent returning deleted records from disk indexes
  185. - Cleared during compaction
  186. ## Performance Characteristics
  187. ### Time Complexity
  188. - **Insert**: O(1) for append, O(log n) for index update
  189. - **Find by primary key (indexed)**: O(log n) binary search
  190. - **Find with query function**: O(n) full scan
  191. - **Update**: O(log n) find + O(1) insert
  192. - **Delete**: O(log n) find + O(1) mark
  193. - **Compact**: O(n) full scan + O(n log n) index rebuild
  194. ### Space Complexity
  195. - Data file grows with inserts (append-only)
  196. - Deleted records remain until compaction
  197. - Index files: O(n) per indexed field
  198. - Delta indexes: O(m) where m = changes since last persist
  199. ### File Size Comparison
  200. MessagePack typically produces **15-20% smaller files** than BSON for the same data:
  201. - More compact integer encoding
  202. - Smaller string overhead
  203. - Efficient array/map encoding
  204. ## Concurrency Model
  205. - **Single-writer, multiple-reader** via file locking
  206. - Writes are serialized through lock file
  207. - Reads can happen concurrently (no locks needed)
  208. - Re-entrant locks allow `withLock()` to wrap compound operations atomically
  209. - Unique indexes enforce constraints under the write lock (no duplicates even with concurrent inserts)
  210. - `refresh()` re-reads `meta.json` before every read so secondary instances see up-to-date state
  211. - Index persistence happens asynchronously but safely
  212. ## Primary Key Types
  213. ### NUMBER (PrimaryKeyType.NUMBER)
  214. - Auto-incremented integer
  215. - Stored in metadata.nextId
  216. - Prefix syntax: `*id`
  217. ### UUID (PrimaryKeyType.UUID)
  218. - Sortable base36 unique ID (9-char timestamp + 3-char random)
  219. - 12-character string format
  220. - Prefix syntax: `@id`
  221. ### STRING (PrimaryKeyType.STRING)
  222. - User-provided string
  223. - No auto-generation
  224. - Default (no prefix)
  225. ## Design Decisions
  226. ### Why Append-Only?
  227. - **Fast writes**: No seeking, just append
  228. - **Crash safety**: Partial writes don't corrupt existing data
  229. - **Simple implementation**: No complex update-in-place logic
  230. ### Why MessagePack?
  231. - **Smaller files**: 15-20% smaller than BSON on average
  232. - **Fast serialization**: Comparable or faster than BSON
  233. - **Wide language support**: Available in many programming languages
  234. - **Simple format**: Easy to implement and debug
  235. - **No external binary dependencies**: Pure JavaScript implementation
  236. ### Why Custom Size Header?
  237. - MessagePack doesn't define record boundaries
  238. - Enables efficient streaming without full deserialization
  239. - Allows skipping deleted records quickly
  240. - Consistent with BSON's approach
  241. ### Why File-Based Locking?
  242. - **Simple**: No external dependencies
  243. - **Cross-process**: Works across multiple Node.js processes
  244. - **Portable**: Works on all platforms
  245. ### Why Binary Search Indexes?
  246. - **Disk-friendly**: Can search large indexes without loading into memory
  247. - **Simple format**: Plain text, easy to debug
  248. - **Fast lookups**: O(log n) for indexed queries
  249. ## MessagePack vs BSON
  250. ### Advantages of MessagePack
  251. - **Smaller files**: 15-20% size reduction
  252. - **Faster reads**: Simpler format, less parsing overhead
  253. - **Pure JavaScript**: No native dependencies
  254. - **Smaller library**: ~15KB vs ~173KB for BSON
  255. ### Advantages of BSON
  256. - **ObjectId type**: Built-in unique identifier type
  257. - **Date precision**: Millisecond timestamps
  258. - **Binary data**: Native binary type
  259. - **MongoDB compatibility**: Direct compatibility with MongoDB
  260. ### When to Choose MPackDB
  261. - File size is a concern
  262. - Pure JavaScript dependencies preferred
  263. - Don't need MongoDB compatibility
  264. - Want faster read performance
  265. ### When to Choose BsonDB
  266. - Need ObjectId primary keys
  267. - MongoDB compatibility desired
  268. - Working with binary data
  269. - Need precise date/time handling
  270. ## Limitations
  271. 1. **Single-writer**: Only one write operation at a time
  272. 2. **No multi-record transactions**: `withLock` provides atomicity for compound operations on a single DB, but not across multiple databases
  273. 3. **No query language**: Must use JavaScript functions for complex queries
  274. 4. **Compaction required**: Deleted records consume space until compaction
  275. 5. **Index overhead**: Each index doubles storage for that field
  276. 6. **No schema validation**: Records can have any structure
  277. ## Future Improvements
  278. - Batch insert operations
  279. - Async compaction (background process)
  280. - Query optimizer for complex filters
  281. - Compression support (MessagePack supports extensions)
  282. - Replication/backup utilities
  283. - Schema validation layer
  284. - Custom MessagePack extension types

Branches

Latest commits

  • b8ffc1a0release 1.0.5caramboleyo
  • d47876a1reimplemented lost features like indexed find and more testscaramboleyo
  • 7f08da9afixed insert ignoring model definitioncaramboleyo
  • 705774a9added flush before findcaramboleyo
  • b4db6391initial commitcaramboleyo