01Embed หรือ Reference
Embed เมื่อข้อมูลมี lifecycle เดียวกัน, ถูกอ่านพร้อมกัน และขนาดไม่โตไม่สิ้นสุด ข้อดีคืออ่านจบในครั้งเดียวและ update ภายใน document เดียวเป็น atomic ส่วน Reference เหมาะกับ many-to-many, entity ที่เปลี่ยนแยกกัน หรือ array ที่โตแบบ unbounded
ที่อยู่ใน policy application, snapshot ของข้อมูลที่ต้อง audit และ item จำนวนจำกัด
customer ที่แชร์หลาย policy, catalog กลาง หรือความสัมพันธ์จำนวนมากและเปลี่ยนบ่อย
02Index และ Query Shape
Compound index ใช้ prefix ของ index จึงต้องเรียง field ให้ตรง query สำคัญ โดยทั่วไป equality fields มาก่อน sort และ range แต่ควรยืนยันด้วยexplain("executionStats") ดู `totalDocsExamined`, `totalKeysExamined` และจำนวนผลลัพธ์
// Query pattern: รายการคำขอล่าสุดของลูกค้า แยกตามสถานะ
db.policyApplications.createIndex({
customerId: 1,
status: 1,
createdAt: -1
})
db.policyApplications.find({
customerId: customerId,
status: "PENDING"
}).sort({ createdAt: -1 }).limit(20).explain("executionStats")ทุก index เพิ่มต้นทุน write และ memory working set หลีกเลี่ยง index ซ้ำซ้อนและ array หลาย field ใน compound index ซึ่งอาจสร้าง multikey entries จำนวนมาก ตรวจ slow query และ index usage ก่อนตัดสินใจเพิ่ม
03Atomicity, Transactions และ Consistency
MongoDB รับประกัน atomicity ระดับ document ดังนั้นการ embed ช่วยลดความจำเป็นของ multi-document transaction หาก workflow ต้องเปลี่ยนหลาย document พร้อมกันสามารถใช้ transaction ได้ แต่มีต้นทุนและไม่ควรใช้เพื่อชดเชย schema ที่ไม่ตรง access pattern
Read concern กำหนดระดับความสอดคล้องของการอ่าน ส่วน write concern กำหนดว่าต้องรอ acknowledgement จาก replica มากเท่าไร การเลือกต้องสัมพันธ์กับ business risk เช่นข้อมูลชำระเงินควรให้ durability สูงกว่าข้อมูล analytics ชั่วคราว
04Replication และ Sharding
Replica set ให้ high availability และ failover ส่วน sharding กระจายข้อมูลและ throughput การเลือก shard key ต้องดู cardinality, distribution และ query targeting key แบบ monotonic อาจทำให้ write ไปรวมที่ shard เดียว ขณะที่ key แบบสุ่มกระจายดีแต่ range query ยากขึ้น
อ่านเอกสารทางการ: MongoDB Embedded Data Models ↗