gitoriaLog in with ident

mpackdb

All repositories: gitoria

ReadmeCodePull requestsReleasesTicketsSettings
Commit705774a9705774a9added flush before findcaramboleyo705774a9/ARCHITECTURE.md

9.9 KB

  1. # MPackDB Architecture
  2. ## Overview
  3. MPackDB is a fast, local, append-only JSON database that uses MessagePack serialization. It's designed for Node.js/Bun applications that need a simple, file-based database with optional indexing capabilities and smaller file sizes compared to BSON.
  4. ## Core Components
  5. ### 1. MPackDB Class (`src/MPackDB.js`)
  6. The main database class that handles all CRUD operations and coordinates between components.
  7. **Key Responsibilities:**
  8. - Database initialization and lifecycle management
  9. - CRUD operations (insert, update, delete, find)
  10. - File locking for concurrent access
  11. - Metadata persistence
  12. - Compaction of deleted records
  13. **Key Properties:**
  14. - `_dataPath`: Path to the `.mpack` data file
  15. - `_dataStream`: Write stream for append-only operations
  16. - `_meta`: Metadata object containing `nextId` and `deleted` offsets
  17. - `_indexManager`: Optional IndexManager instance for indexed queries
  18. - `_primaryKey`: Name of the primary key field
  19. - `_primaryKeyType`: Type of primary key (NUMBER, UUID, STRING)
  20. ### 2. IndexManager Class (`src/IndexManager.js`)
  21. Manages binary search indexes for fast lookups on indexed fields.
  22. **Key Responsibilities:**
  23. - Building and maintaining indexes from data files
  24. - Binary search on disk-based index files
  25. - Delta indexes (in-memory changes not yet persisted)
  26. - Tombstone tracking for deleted records
  27. - Auto-persistence of indexes
  28. **Index File Format:**
  29. ```
  30. key,offset,length
  31. key,offset,length
  32. ...
  33. ```
  34. Each line represents an index entry where:
  35. - `key`: The indexed field value
  36. - `offset`: Byte offset in the data file
  37. - `length`: Length of the record in bytes
  38. **Index Types:**
  39. - **NUMERIC**: Numeric comparison for sorting/searching
  40. - **LEXICAL**: String comparison for sorting/searching
  41. ### 3. Cursor Class (`src/Cursor.js`)
  42. Provides an async iterable interface for query results.
  43. **Key Responsibilities:**
  44. - Lazy evaluation of queries
  45. - Support for different iteration modes (record, offset, mixed, raw)
  46. - Integration with IndexManager for indexed queries
  47. - Filtering with query functions
  48. ### 4. MessagePack Utilities (`src/mpack.js`)
  49. Wrapper around `msgpackr` with custom utilities.
  50. **Key Exports:**
  51. - `serialize()`: Encode JavaScript objects to MessagePack binary
  52. - `deserialize()`: Decode MessagePack binary to JavaScript objects
  53. - `uuid()`: Generate sortable 12-char base36 unique IDs (9 timestamp + 3 random)
  54. - `PrimaryKeyType`: Enum for primary key types
  55. - `IndexType`: Enum for index types
  56. ## Data Flow
  57. ### Insert Operation
  58. ```
  59. 1. User calls db.insert(record)
  60. 2. Acquire file lock
  61. 3. Auto-generate primary key if needed (numeric/UUID)
  62. 4. Serialize record to MessagePack
  63. 5. Prepend 4-byte size header
  64. 6. Append to data file via write stream
  65. 7. Add entry to IndexManager (if indexes enabled)
  66. 8. Persist metadata (if nextId changed)
  67. 9. Release file lock
  68. 10. Return primary key or full record
  69. ```
  70. ### Find Operation (Indexed)
  71. ```
  72. 1. User calls db.find(primaryKeyValue)
  73. 2. Cursor created with query function
  74. 3. IndexManager performs binary search on index file
  75. 4. Check delta indexes for recent changes
  76. 5. Check tombstones for deleted records
  77. 6. Read record from data file at found offset
  78. 7. Yield record to user
  79. ```
  80. ### Find Operation (Non-Indexed)
  81. ```
  82. 1. User calls db.find(queryFn)
  83. 2. Cursor created with query function
  84. 3. Stream through entire data file
  85. 4. Read 4-byte size header
  86. 5. Read MessagePack data based on size
  87. 6. Deserialize each record
  88. 7. Apply query function filter
  89. 8. Skip deleted records (check metadata.deleted)
  90. 9. Yield matching records to user
  91. ```
  92. ### Update Operation
  93. ```
  94. 1. User calls db.update(query, newData)
  95. 2. Find matching records (using find)
  96. 3. For each match:
  97. a. Mark old record as deleted
  98. b. Insert new record with same primary key
  99. 4. Persist metadata with new deleted offsets
  100. ```
  101. ### Delete Operation
  102. ```
  103. 1. User calls db.delete(query)
  104. 2. Find matching records
  105. 3. Add offsets to metadata.deleted array
  106. 4. Remove from indexes (if enabled)
  107. 5. Persist metadata
  108. ```
  109. ### Compact Operation
  110. ```
  111. 1. Acquire file lock
  112. 2. Create temporary data file
  113. 3. Stream through all records
  114. 4. Write only non-deleted records to temp file
  115. 5. Atomically rename temp file to replace original
  116. 6. Clear metadata.deleted array
  117. 7. Rebuild all indexes from new file
  118. 8. Release file lock
  119. ```
  120. ## File Structure
  121. ```
  122. data/
  123. ├── users.mpack # Main data file (MessagePack records)
  124. ├── users.meta.json # Metadata (nextId, deleted offsets)
  125. ├── users.lock # Lock file (contains PID)
  126. ├── users.id.txt # Index file for 'id' field
  127. ├── users.email.txt # Index file for 'email' field
  128. └── ...
  129. ```
  130. ## Serialization Format
  131. MPackDB uses MessagePack format from the `msgpackr` npm package with a custom size header:
  132. ```
  133. [4 bytes: size][MessagePack data]
  134. ```
  135. Each record is prefixed with a 4-byte little-endian integer indicating the size of the MessagePack data (not including the size prefix itself).
  136. **Why the size header?**
  137. - MessagePack doesn't include record boundaries in the format
  138. - The size header allows streaming reads without parsing the entire file
  139. - Enables skipping deleted records efficiently
  140. - Matches the pattern used in BSON for consistency
  141. ## Locking Mechanism
  142. MPackDB uses file-based locking to prevent concurrent writes:
  143. 1. Before any write operation, create `{dbPath}.lock` file with `wx` flag (exclusive)
  144. 2. Write current process PID to lock file
  145. 3. If lock exists, wait 100ms and retry
  146. 4. After operation completes, delete lock file
  147. This ensures only one process can write at a time while allowing multiple readers.
  148. ## Index Persistence Strategy
  149. Indexes use a two-tier approach:
  150. ### Disk Indexes
  151. - Sorted index files on disk
  152. - Binary searchable for O(log n) lookups
  153. - Rebuilt during compaction
  154. ### Delta Indexes (In-Memory)
  155. - Track changes since last persistence
  156. - Checked before disk indexes
  157. - Auto-persisted based on:
  158. - Time interval (default: 60 seconds)
  159. - Change threshold (default: 1000 operations)
  160. ### Tombstones
  161. - Track deleted records in memory
  162. - Prevent returning deleted records from disk indexes
  163. - Cleared during compaction
  164. ## Performance Characteristics
  165. ### Time Complexity
  166. - **Insert**: O(1) for append, O(log n) for index update
  167. - **Find by primary key (indexed)**: O(log n) binary search
  168. - **Find with query function**: O(n) full scan
  169. - **Update**: O(log n) find + O(1) insert
  170. - **Delete**: O(log n) find + O(1) mark
  171. - **Compact**: O(n) full scan + O(n log n) index rebuild
  172. ### Space Complexity
  173. - Data file grows with inserts (append-only)
  174. - Deleted records remain until compaction
  175. - Index files: O(n) per indexed field
  176. - Delta indexes: O(m) where m = changes since last persist
  177. ### File Size Comparison
  178. MessagePack typically produces **15-20% smaller files** than BSON for the same data:
  179. - More compact integer encoding
  180. - Smaller string overhead
  181. - Efficient array/map encoding
  182. ## Concurrency Model
  183. - **Single-writer, multiple-reader** via file locking
  184. - Writes are serialized through lock file
  185. - Reads can happen concurrently (no locks needed)
  186. - Index persistence happens asynchronously but safely
  187. ## Primary Key Types
  188. ### NUMBER (PrimaryKeyType.NUMBER)
  189. - Auto-incremented integer
  190. - Stored in metadata.nextId
  191. - Prefix syntax: `*id`
  192. ### UUID (PrimaryKeyType.UUID)
  193. - Sortable base36 unique ID (9-char timestamp + 3-char random)
  194. - 12-character string format
  195. - Prefix syntax: `@id`
  196. ### STRING (PrimaryKeyType.STRING)
  197. - User-provided string
  198. - No auto-generation
  199. - Default (no prefix)
  200. ## Design Decisions
  201. ### Why Append-Only?
  202. - **Fast writes**: No seeking, just append
  203. - **Crash safety**: Partial writes don't corrupt existing data
  204. - **Simple implementation**: No complex update-in-place logic
  205. ### Why MessagePack?
  206. - **Smaller files**: 15-20% smaller than BSON on average
  207. - **Fast serialization**: Comparable or faster than BSON
  208. - **Wide language support**: Available in many programming languages
  209. - **Simple format**: Easy to implement and debug
  210. - **No external binary dependencies**: Pure JavaScript implementation
  211. ### Why Custom Size Header?
  212. - MessagePack doesn't define record boundaries
  213. - Enables efficient streaming without full deserialization
  214. - Allows skipping deleted records quickly
  215. - Consistent with BSON's approach
  216. ### Why File-Based Locking?
  217. - **Simple**: No external dependencies
  218. - **Cross-process**: Works across multiple Node.js processes
  219. - **Portable**: Works on all platforms
  220. ### Why Binary Search Indexes?
  221. - **Disk-friendly**: Can search large indexes without loading into memory
  222. - **Simple format**: Plain text, easy to debug
  223. - **Fast lookups**: O(log n) for indexed queries
  224. ## MessagePack vs BSON
  225. ### Advantages of MessagePack
  226. - **Smaller files**: 15-20% size reduction
  227. - **Faster reads**: Simpler format, less parsing overhead
  228. - **Pure JavaScript**: No native dependencies
  229. - **Smaller library**: ~15KB vs ~173KB for BSON
  230. ### Advantages of BSON
  231. - **ObjectId type**: Built-in unique identifier type
  232. - **Date precision**: Millisecond timestamps
  233. - **Binary data**: Native binary type
  234. - **MongoDB compatibility**: Direct compatibility with MongoDB
  235. ### When to Choose MPackDB
  236. - File size is a concern
  237. - Pure JavaScript dependencies preferred
  238. - Don't need MongoDB compatibility
  239. - Want faster read performance
  240. ### When to Choose BsonDB
  241. - Need ObjectId primary keys
  242. - MongoDB compatibility desired
  243. - Working with binary data
  244. - Need precise date/time handling
  245. ## Limitations
  246. 1. **Single-writer**: Only one write operation at a time
  247. 2. **No transactions**: Operations are not atomic across multiple records
  248. 3. **No query language**: Must use JavaScript functions for complex queries
  249. 4. **Compaction required**: Deleted records consume space until compaction
  250. 5. **Index overhead**: Each index doubles storage for that field
  251. 6. **No schema validation**: Records can have any structure
  252. 7. **No ObjectId type**: Must use UUIDs or numeric IDs
  253. ## Future Improvements
  254. - Batch insert operations
  255. - Async compaction (background process)
  256. - Query optimizer for complex filters
  257. - Compression support (MessagePack supports extensions)
  258. - Replication/backup utilities
  259. - Schema validation layer
  260. - Custom MessagePack extension types

Branches

Latest commits

  • 705774a9added flush before findcaramboleyo
  • b4db6391initial commitcaramboleyo