gitoriaLog in with ident

mpackdb

All repositories: gitoria

ReadmeCodePull requestsReleasesTicketsSettings
Commitcde73eb4cde73eb4release 1.0.6caramboleyocde73eb4/ARCHITECTURE.md

15.2 KB

  1. # MPackDB Architecture
  2. ## Overview
  3. MPackDB is a fast, local, append-only JSON database that uses MessagePack serialization. It's designed for Node.js/Bun applications that need a simple, file-based database with optional indexing capabilities and smaller file sizes compared to BSON.
  4. ## Core Components
  5. ### 1. MPackDB Class (`src/MPackDB.js`)
  6. The main database class that handles all CRUD operations and coordinates between components.
  7. **Key Responsibilities:**
  8. - Database initialization and lifecycle management
  9. - CRUD operations (insert, update, delete, find)
  10. - File locking for concurrent access
  11. - Metadata persistence
  12. - Compaction of deleted records
  13. **Key Properties:**
  14. - `_dataPath`: Path to the `.mpack` data file
  15. - `_dataStream`: Write stream for append-only operations
  16. - `_meta`: Metadata object containing `nextId` and `deleted` offsets
  17. - `_indexManager`: Optional IndexManager instance for indexed queries
  18. - `_primaryKey`: Name of the primary key field
  19. - `_primaryKeyType`: Type of primary key (NUMBER, UUID, STRING)
  20. ### 2. IndexManager Class (`src/IndexManager.js`)
  21. Manages binary search indexes for fast lookups on indexed fields.
  22. **Key Responsibilities:**
  23. - Building and maintaining indexes from data files
  24. - Binary search on disk-based index files
  25. - Delta indexes (in-memory changes not yet persisted)
  26. - Tombstone tracking for deleted records
  27. - Auto-persistence of indexes
  28. **Index File Format:**
  29. ```
  30. key,offset,length
  31. key,offset,length
  32. ...
  33. ```
  34. Each line represents an index entry where:
  35. - `key`: The indexed field value
  36. - `offset`: Byte offset in the data file
  37. - `length`: Length of the record in bytes
  38. **Index Types:**
  39. - **NUMERIC**: Numeric comparison for sorting/searching
  40. - **LEXICAL**: String comparison for sorting/searching
  41. - **UNIQUE**: Any index can be marked unique (`!` prefix) to reject duplicates on insert
  42. ### 3. Cursor Class (`src/Cursor.js`)
  43. Provides an async iterable interface for query results.
  44. **Key Responsibilities:**
  45. - Lazy evaluation of queries
  46. - Support for different iteration modes (record, offset, mixed, raw)
  47. - Integration with IndexManager for indexed queries
  48. - Filtering with query functions
  49. ### 4. MessagePack Utilities (`src/mpack.js`)
  50. Wrapper around `msgpackr` with custom utilities.
  51. **Key Exports:**
  52. - `serialize()`: Encode JavaScript objects to MessagePack binary
  53. - `deserialize()`: Decode MessagePack binary to JavaScript objects
  54. - `uuid()`: Generate sortable 12-char base36 unique IDs (9 timestamp + 3 random)
  55. - `PrimaryKeyType`: Enum for primary key types
  56. - `IndexType`: Enum for index types
  57. ## Data Flow
  58. ### Insert Operation
  59. ```
  60. 1. User calls db.insert(record)
  61. 2. Acquire file lock
  62. 3. Auto-generate primary key if needed (numeric/UUID)
  63. 4. Check unique index constraints (skip auto-generated PKs)
  64. 5. Serialize record to MessagePack
  65. 6. Prepend 4-byte size header
  66. 7. Append to data file via write stream
  67. 8. Add entry to IndexManager (if indexes enabled)
  68. 9. Persist metadata (if nextId changed)
  69. 10. Release file lock
  70. 11. Return primary key or full record
  71. ```
  72. ### Find Operation (Indexed)
  73. ```
  74. 1. User calls db.find(primaryKeyValue)
  75. 2. Cursor created with query function
  76. 3. IndexManager performs binary search on index file
  77. 4. Check delta indexes for recent changes
  78. 5. Check tombstones for deleted records
  79. 6. Read record from data file at found offset
  80. 7. Yield record to user
  81. ```
  82. ### Find Operation (Non-Indexed)
  83. ```
  84. 1. User calls db.find(queryFn)
  85. 2. Cursor created with query function
  86. 3. Stream through entire data file
  87. 4. Read 4-byte size header
  88. 5. Read MessagePack data based on size
  89. 6. Deserialize each record
  90. 7. Apply query function filter
  91. 8. Skip deleted records (check metadata.deleted)
  92. 9. Yield matching records to user
  93. ```
  94. ### Find with Index Hints
  95. ```
  96. 1. User calls db.find(filterFn, { index: { field, from, to, direction } })
  97. or db.find(filterFn, { index: [hint1, hint2, ...] })
  98. 2. For each index hint, collect offset sets:
  99. - value hint: exact match via get() → offset set
  100. - range hint: entries(from, to, direction) → offset set
  101. 3. If multiple hints: intersect all offset sets (smallest first)
  102. 4. Stream records from the resulting offsets
  103. 5. Apply filter function on deserialized records
  104. 6. Yield matching records
  105. ```
  106. ### Bounding Box Query
  107. ```
  108. 1. User calls db.boundingBox(corners, filter)
  109. 2. Extract min/max for x and y from corners (order irrelevant)
  110. 3. Delegates to find(filter, { index: [
  111. { field: 'x', from: minX, to: maxX },
  112. { field: 'y', from: minY, to: maxY }
  113. ]})
  114. 4. Intersection + streaming as above
  115. ```
  116. ### Update Operation
  117. ```
  118. 1. User calls db.update(query, newData, { index })
  119. 2. Internally calls delete(query, { index, callback })
  120. 3. For each match (using find with optional index hints):
  121. a. Mark old record as deleted
  122. b. Callback inserts new record with same primary key
  123. 4. Persist metadata with new deleted offsets
  124. ```
  125. ### Delete Operation
  126. ```
  127. 1. User calls db.delete(query, { index })
  128. 2. Find matching records (uses index hints if provided, otherwise full scan)
  129. 3. Add offsets to metadata.deleted array
  130. 4. Remove from indexes (if enabled)
  131. 5. Persist metadata
  132. ```
  133. ### Compact Operation
  134. ```
  135. 1. Acquire file lock
  136. 2. Create temporary data file
  137. 3. Stream through all records
  138. 4. Write only non-deleted records to temp file
  139. 5. Atomically rename temp file to replace original
  140. 6. Clear metadata.deleted array
  141. 7. Rebuild all indexes from new file
  142. 8. Release file lock
  143. ```
  144. ## File Structure
  145. ```
  146. data/
  147. ├── users.mpack # Main data file (MessagePack records)
  148. ├── users.meta.json # Metadata (nextId, deleted offsets)
  149. ├── users.lock # Lock file (contains PID)
  150. ├── users.id.txt # Index file for 'id' field
  151. ├── users.email.txt # Index file for 'email' field
  152. └── ...
  153. ```
  154. ## Serialization Format
  155. MPackDB uses MessagePack format from the `msgpackr` npm package with a custom size header:
  156. ```
  157. [4 bytes: size][MessagePack data]
  158. ```
  159. Each record is prefixed with a 4-byte little-endian integer indicating the size of the MessagePack data (not including the size prefix itself).
  160. **Why the size header?**
  161. - MessagePack doesn't include record boundaries in the format
  162. - The size header allows streaming reads without parsing the entire file
  163. - Enables skipping deleted records efficiently
  164. - Matches the pattern used in BSON for consistency
  165. ## Locking Mechanism
  166. MPackDB uses file-based locking to prevent concurrent writes:
  167. 1. Before any write operation, create `{dbPath}.lock` file with `wx` flag (exclusive)
  168. 2. Write current process PID to lock file
  169. 3. If lock exists, wait 100ms and retry
  170. 4. After operation completes, delete lock file
  171. The lock is **re-entrant**: if a lock is already held (e.g. inside `withLock()`), nested operations (`insert`, `delete`, etc.) increment a depth counter instead of acquiring a new file lock. The file lock is only released when the outermost holder finishes.
  172. ### withLock for Compound Operations
  173. `db.withLock(callback)` acquires the lock for the duration of the callback. All DB operations inside the callback reuse the same lock. This makes compound operations like find-then-insert atomic:
  174. ```javascript
  175. await db.withLock(async () => {
  176. const [user] = await db.find(u => u.email === email);
  177. if (!user) await db.insert({ email });
  178. });
  179. ```
  180. This ensures only one process can write at a time while allowing multiple readers.
  181. ## Index Persistence Strategy
  182. Indexes use a two-tier approach:
  183. ### Disk Indexes
  184. - Sorted index files on disk
  185. - Binary searchable for O(log n) lookups
  186. - Rebuilt during compaction
  187. ### Delta Indexes (In-Memory)
  188. - Track changes since last persistence
  189. - Checked before disk indexes
  190. - Auto-persisted based on:
  191. - Time interval (default: 60 seconds)
  192. - Change threshold (default: 1000 operations)
  193. ### Tombstones
  194. - Track deleted records in memory
  195. - Prevent returning deleted records from disk indexes
  196. - Cleared during compaction
  197. ## Performance Characteristics
  198. ### Time Complexity
  199. - **Insert**: O(1) for append, O(log n) for index update
  200. - **Find by primary key (indexed)**: O(log n) binary search
  201. - **Find with query function**: O(n) full scan
  202. - **Find with single index hint**: O(log n) seek + O(k) scan where k = entries in range
  203. - **Find with index intersection**: O(k1 + k2 + ... + min(k)) for collecting + intersecting offset sets, then O(m) disk reads where m = intersection size
  204. - **Bounding box**: Same as index intersection with 2 range indexes
  205. - **Update**: O(find) + O(1) insert
  206. - **Delete**: O(find) + O(1) mark
  207. - **Compact**: O(n) full scan + O(n log n) index rebuild
  208. ### Space Complexity
  209. - Data file grows with inserts (append-only)
  210. - Deleted records remain until compaction
  211. - Index files: O(n) per indexed field
  212. - Delta indexes: O(m) where m = changes since last persist
  213. ### File Size Comparison
  214. MessagePack typically produces **15-20% smaller files** than BSON for the same data:
  215. - More compact integer encoding
  216. - Smaller string overhead
  217. - Efficient array/map encoding
  218. ## Concurrency Model
  219. - **Single-writer, multiple-reader** via file locking
  220. - Writes are serialized through lock file
  221. - Reads can happen concurrently (no locks needed)
  222. - Re-entrant locks allow `withLock()` to wrap compound operations atomically
  223. - Unique indexes enforce constraints under the write lock (no duplicates even with concurrent inserts)
  224. - `refresh()` re-reads `meta.json` before every read so secondary instances see up-to-date state
  225. - Index persistence happens asynchronously but safely
  226. ## Primary Key Types
  227. ### NUMBER (PrimaryKeyType.NUMBER)
  228. - Auto-incremented integer
  229. - Stored in metadata.nextId
  230. - Prefix syntax: `*id`
  231. ### UUID (PrimaryKeyType.UUID)
  232. - Sortable base36 unique ID (9-char timestamp + 3-char random)
  233. - 12-character string format
  234. - Prefix syntax: `@id`
  235. ### STRING (PrimaryKeyType.STRING)
  236. - User-provided string
  237. - No auto-generation
  238. - Default (no prefix)
  239. ## Design Decisions
  240. ### Why Append-Only?
  241. - **Fast writes**: No seeking, just append
  242. - **Crash safety**: Partial writes don't corrupt existing data
  243. - **Simple implementation**: No complex update-in-place logic
  244. ### Why MessagePack?
  245. - **Smaller files**: 15-20% smaller than BSON on average
  246. - **Fast serialization**: Comparable or faster than BSON
  247. - **Wide language support**: Available in many programming languages
  248. - **Simple format**: Easy to implement and debug
  249. - **No external binary dependencies**: Pure JavaScript implementation
  250. ### Why Custom Size Header?
  251. - MessagePack doesn't define record boundaries
  252. - Enables efficient streaming without full deserialization
  253. - Allows skipping deleted records quickly
  254. - Consistent with BSON's approach
  255. ### Why File-Based Locking?
  256. - **Simple**: No external dependencies
  257. - **Cross-process**: Works across multiple Node.js processes
  258. - **Portable**: Works on all platforms
  259. ### Why Binary Search Indexes?
  260. - **Disk-friendly**: Can search large indexes without loading into memory
  261. - **Simple format**: Plain text, easy to debug
  262. - **Fast lookups**: O(log n) for indexed queries
  263. ## MessagePack vs BSON
  264. ### Advantages of MessagePack
  265. - **Smaller files**: 15-20% size reduction
  266. - **Faster reads**: Simpler format, less parsing overhead
  267. - **Pure JavaScript**: No native dependencies
  268. - **Smaller library**: ~15KB vs ~173KB for BSON
  269. ### Advantages of BSON
  270. - **ObjectId type**: Built-in unique identifier type
  271. - **Date precision**: Millisecond timestamps
  272. - **Binary data**: Native binary type
  273. - **MongoDB compatibility**: Direct compatibility with MongoDB
  274. ### When to Choose MPackDB
  275. - File size is a concern
  276. - Pure JavaScript dependencies preferred
  277. - Don't need MongoDB compatibility
  278. - Want faster read performance
  279. ### When to Choose BsonDB
  280. - Need ObjectId primary keys
  281. - MongoDB compatibility desired
  282. - Working with binary data
  283. - Need precise date/time handling
  284. ## Limitations
  285. 1. **Single-writer**: Only one write operation at a time
  286. 2. **No multi-record transactions**: `withLock` provides atomicity for compound operations on a single DB, but not across multiple databases
  287. 3. **No query language**: Must use JavaScript functions for complex queries
  288. 4. **Compaction required**: Deleted records consume space until compaction
  289. 5. **Index overhead**: Each index doubles storage for that field
  290. 6. **No schema validation**: Records can have any structure
  291. ## Index Query Design Decisions
  292. ### Why `to` but not `limit` or `filter`
  293. The `entries()` generator supports `from`, `to`, and `direction` but intentionally has no `limit` or `filter` parameters.
  294. **Why `to` is needed:** When collecting offset sets for index intersection, all matching offsets are loaded into a Map. Without `to`, a range scan like `{ field: 'x', from: 0 }` would collect every entry from 0 to the end of the index — potentially millions of offsets when only a small range is needed. `to` stops the scan at the upper bound, keeping the offset set small.
  295. **Why `limit` was removed:** `limit` caps the number of results, but the caller doesn't know the right count upfront. Since results are streamed via async generators, the caller can `break` out of the loop at any time, which terminates the generator immediately. This is more flexible than a fixed limit and works naturally with JavaScript's `for await` syntax.
  296. **Why `filter` was removed:** With streaming, the caller filters in the loop body after deserialization. A built-in filter would only save one `if` statement per iteration while adding API complexity. The filter function in `find()` handles this at the right layer — after records are fetched from disk.
  297. ### Why index intersection uses pre-collected offset sets
  298. Index intersection collects full offset sets from each index hint, then intersects them before reading any records from disk. This trades memory (storing offset integers) for disk I/O (avoiding unnecessary record reads).
  299. **Trade-off:** For a query like `x: 0..100 AND status: 'active'` on 10 million records, the x-index might yield 500k offsets and the status-index 50k offsets. Intersecting these sets (~4MB of integers in memory) reduces disk reads from 500k to maybe 10k — a massive I/O saving.
  300. **Why not stream-and-check:** An alternative would be to stream the primary index and check each candidate against secondary indexes on the fly (O(log n) lookup per candidate). This uses less memory but requires a disk seek per candidate. For large datasets where the secondary index is highly selective, pre-collected intersection wins. For small datasets, the difference is negligible.
  301. **Why the caller picks the primary index:** The engine cannot automatically pick the most selective index without scanning all of them first. Instead, the caller specifies which indexes to use via the `index` option. For streaming without intersection (single index), the caller controls the scan with `from`, `to`, `direction`, and `break`.
  302. ### Bounding box vs geospatial
  303. `boundingBox()` operates on flat 2D coordinates — simple min/max comparisons on x and y axes. This is sufficient for game grids, tile maps, and any coordinate system where geometry is planar. Geospatial indexing (S2 cells, geohash) is only needed for spherical geometry (lat/lng on Earth) where "straight lines" are actually great-circle arcs and distances warp with latitude. Both approaches use the same fundamental pattern internally: coarse index scan → exact geometric filter.
  304. ## Future Improvements
  305. - Batch insert operations
  306. - Async compaction (background process)
  307. - Query optimizer for complex filters
  308. - Compression support (MessagePack supports extensions)
  309. - Replication/backup utilities
  310. - Schema validation layer
  311. - Custom MessagePack extension types

Branches

Latest commits

  • cde73eb4release 1.0.6caramboleyo
  • d01dda02add index hints, intersection, boundingBox; remove findByIndexcaramboleyo
  • b8ffc1a0release 1.0.5caramboleyo
  • d47876a1reimplemented lost features like indexed find and more testscaramboleyo
  • 7f08da9afixed insert ignoring model definitioncaramboleyo
  • 705774a9added flush before findcaramboleyo
  • b4db6391initial commitcaramboleyo