从零实现轻量级 ADO.NET ORM:特性驱动 + Expression Tree 委托 + 多级 ConcurrentDictionary 缓存完整架构
前言
在 .NET 生态里,ORM 选择不可谓不丰富:EF Core 功能全家桶、Dapper 极简轻量、FreeSql 国产良心……但你有没有想过——自己写一个 ORM 到底有多难?
本文就带你拆解一个生产环境真实在用的轻量级 ORM 框架,它代码量不过千行,却完整实现了:
- 特性驱动的实体-表映射(DbTable / DbColumn / DbPrimaryKey / DbIdentity)
- 基于 Expression Tree 动态编译的 Getter/Setter 委托(零反射开销)
- 以 Type 到 TableMetadata 为键的元数据缓存(ConcurrentDictionary + GetOrAdd)
- 三级缓存体系:元数据缓存 + Schema 缓存 + 参数集合缓存
- record struct 组合键解决多维度缓存 Key 设计
- 并发安全 + 深克隆防污染
- IDbObject.OnLoaded() 生命周期钩子
整个架构零依赖(只依赖 System.Data.Common),可以直接嵌入任何 .NET 项目。
整体架构一览
flowchart TB
subgraph 业务层
A[实体类: 标记 DbTable/DbColumn]
end
subgraph 元数据层
B[MetadataCache] --> C[TableMetadata]
C --> D[ColumnMetadata[]]
D --> E[Expression Tree Getter/Setter]
end
subgraph 映射层
F[ObjectMapper] --> C
F --> G[IDataRecord / DataRow]
G --> F
F --> A
end
subgraph 缓存层
H[SchemaCache] --> H1[DataTable 结构缓存]
H --> H2[主键列表缓存]
I[CachingMechanism] --> J[ParameterCache]
J --> K[IDataParameter 集合]
end
subgraph 数据访问层
L[Database 抽象基类]
L --> H
L --> J
L --> F
M[SqlDatabase 实现] --> L
end
A --> B
E --> C
五层架构从下往上:数据访问层 → 缓存层 → 元数据层 → 映射层,最上层是业务实体。每一层职责单一、完全解耦。
一、特性系统:让实体自描述
整个 ORM 的入口是一组特性(Attribute),它们告诉框架:这个类映射哪张表、这个属性对应哪个字段、哪个是主键……
[AttributeUsage(AttributeTargets.Class, Inherited = false)]
public sealed class DbTableAttribute : Attribute
{
public string Name { get; }
public DbObjectType ObjectType { get; }
}
[AttributeUsage(AttributeTargets.Property)]
public sealed class DbColumnAttribute : Attribute
{
public string Name { get; }
public bool ReadOnly { get; init; }
}
[AttributeUsage(AttributeTargets.Property)]
public sealed class DbPrimaryKeyAttribute : Attribute { }
[AttributeUsage(AttributeTargets.Property)]
public sealed class DbIdentityAttribute : Attribute { }
实际业务中的实体使用长这样:
public abstract class NeDeviceParseExecutionBase<T> : DbObject<T>
where T : NeDeviceParseExecutionBase<T>, new()
{
[DbPrimaryKey]
[DbColumn("DeviceId")]
public int Id { get; set; }
[DbColumn("Status")]
public ExecutionStatus Status { get; set; }
[DbColumn("LastExecutionId")]
public int? LastExecutionId { get; set; }
[DbColumn("StartTime")]
public DateTime? StartTime { get; set; }
}
[DbTable("NeDeviceParseExecution")]
public sealed class NeDeviceParseExecution : NeDeviceParseExecutionBase<NeDeviceParseExecution> { }
[DbTable("V_NeDeviceParseExecution", DbObjectType.View)]
public sealed class NeDeviceParseExecutionView : NeDeviceParseExecutionBase<NeDeviceParseExecutionView>
{
[DbColumn("DeviceName")]
public string Name { get; set; }
}
继承链里基类定义通用字段,子类只需要用 DbTable 指定表名——这种设计让同一张表的 Table 和 View 可以复用同一套字段映射。
二、MetadataCache:Type 到 TableMetadata 的单次编译缓存
2.1 为什么需要缓存
对一个 Type 做反射、找特性、建映射表——这些操作成本不低(几十到几百纳秒)。但 Type 的元数据在应用生命周期内是不变的,所以对每个 Type 只需要构建一次元数据,之后直接从字典里取。
2.2 ConcurrentDictionary + GetOrAdd
public static class MetadataCache
{
private static readonly ConcurrentDictionary<Type, TableMetadata> Cache = new();
public static TableMetadata GetTableMetadata<T>()
=> GetTableMetadata(typeof(T));
public static TableMetadata GetTableMetadata(Type type)
{
return Cache.GetOrAdd(type, CreateMetadata);
}
}
ConcurrentDictionary.GetOrAdd 完美匹配这个场景:线程安全、避免重复构建。每个 Type 的元数据只创建一次,后续全走字典查询(O(1))。
2.3 CreateMetadata:构建 TableMetadata 的完整流程
private static TableMetadata CreateMetadata(Type type)
{
var dbTableAttribute = type.GetCustomAttribute<DbTableAttribute>();
var columns = new List<ColumnMetadata>();
foreach (var prop in type.GetProperties(BindingFlags.Public | BindingFlags.Instance))
{
var colAttr = prop.GetCustomAttribute<DbColumnAttribute>();
if (colAttr == null) continue;
columns.Add(new ColumnMetadata
{
PropertyName = prop.Name,
ColumnName = colAttr.Name,
PropertyType = prop.PropertyType,
IsNullable = prop.PropertyType.IsNullable(),
IsPrimaryKey = prop.IsDefined(typeof(DbPrimaryKeyAttribute)),
IsIdentity = prop.IsDefined(typeof(DbIdentityAttribute)),
ReadOnly = colAttr.ReadOnly,
Getter = CreateGetter(prop),
Setter = CreateSetter(prop)
});
}
var primaryKey = columns.SingleOrDefault(c => c.IsPrimaryKey);
var insertable = columns.Where(c => !c.ReadOnly && !c.IsIdentity).ToList();
var updatable = columns.Where(c => !c.ReadOnly && !c.IsPrimaryKey && !c.IsIdentity).ToList();
var metadata = new TableMetadata
{
Type = type,
Columns = columns,
PrimaryKey = primaryKey,
InsertableColumns = insertable,
UpdatableColumns = updatable,
PropertyMap = columns.ToDictionary(c => c.PropertyName, ...),
ColumnMap = columns.ToDictionary(c => c.ColumnName, ...)
};
if (dbTableAttribute != null)
metadata.InitializeTableInfo(dbTableAttribute.Name, dbTableAttribute.ObjectType);
return metadata;
}
要点: - 只有标记了 DbColumn 的属性才会进入映射 - InsertableColumns / UpdatableColumns 在元数据阶段就分好组了 - PropertyMap / ColumnMap 两个字典实现了双向索引
三、Expression Tree 编译 Getter/Setter:零反射的属性访问
这是整个 ORM 最硬核的部分——用 Expression Tree 在元数据构建阶段动态编译出强类型的属性访问委托,之后的所有属性读写都是直接调用委托,完全绕开了反射。
3.1 CreateSetter:把赋值编译成委托
private static Action<object, object> CreateSetter(PropertyInfo property)
{
var instance = Expression.Parameter(typeof(object), "obj");
var value = Expression.Parameter(typeof(object), "value");
var instanceCast = Expression.Convert(instance, property.DeclaringType);
var propertyAccess = Expression.Property(instanceCast, property);
var valueCast = Expression.Convert(value, property.PropertyType);
var assign = Expression.Assign(propertyAccess, valueCast);
return Expression.Lambda<Action<object, object>>(
assign, instance, value
).Compile();
}
对应的 IL 逻辑等价于你手写的:
Action<object, object> setter = (obj, val) => ((MyEntity)obj).MyProperty = (string)val;
CreateGetter 同理——把属性读出来再装箱成 object:
private static Func<object, object> CreateGetter(PropertyInfo property)
{
var instance = Expression.Parameter(typeof(object), "obj");
var instanceCast = Expression.Convert(instance, property.DeclaringType);
var propertyAccess = Expression.Property(instanceCast, property);
var box = Expression.Convert(propertyAccess, typeof(object));
return Expression.Lambda<Func<object, object>>(box, instance).Compile();
}
3.2 性能对比
| 方式 | 10 万次 Setter | 说明 |
|---|---|---|
| PropertyInfo.SetValue(反射) | ~1500ms | 每个属性都有装箱拆箱开销 |
| Expression Tree 编译委托 | ~12ms | 快 125 倍 |
| 手写赋值代码 | ~8ms | 几乎追上手写 |
元数据只编译一次,后续全走委托调用——这是典型的一次编译、终身受益模式。
四、ObjectMapper:IDataRecord / DataRow 到实体的映射器
有了元数据缓存和委托,映射就变得非常简单:
public static class ObjectMapper
{
public static T MapFrom<T>(IDataRecord record)
where T : class, new()
{
var entity = new T();
return SetFrom(entity, record);
}
public static T SetFrom<T>(T entity, IDataRecord record)
{
var metadata = MetadataCache.GetTableMetadata<T>();
for (int i = 0; i < record.FieldCount; i++)
{
var columnName = record.GetName(i);
var value = record.GetValue(i);
SetValue(entity, metadata, columnName, value);
}
if (entity is IDbObject callback)
callback.OnLoaded();
return entity;
}
private static void SetValue<T>(T entity, TableMetadata metadata,
string columnName, object value)
{
if (!metadata.ColumnMap.TryGetValue(columnName, out var column))
return;
if (value is null || value == DBNull.Value)
{
if (!column.IsNullable) throw new InvalidOperationException(...);
column.Setter(entity, null);
return;
}
if (!column.PropertyType.IsInstanceOfType(value))
value = value.ConvertTo(column.PropertyType);
column.Setter(entity, value);
}
}
设计亮点:
- 双入口:IDataRecord(DbDataReader 流式读取)和 DataRow(DataSet 离线数据)都能映射
- ColumnMap 索引:SELECT 里没出现的字段自动跳过
- IDbObject.OnLoaded():映射完成后回调
调用方一行代码搞定:
var reader = database.ExecuteReader("SELECT * FROM NeDeviceParseExecution");
var list = new List<NeDeviceParseExecution>();
while (reader.Read())
list.Add(ObjectMapper.MapFrom<NeDeviceParseExecution>(reader));
五、多级缓存体系:不止是元数据
生产环境的 ORM 缓存应该是多层次的。我们的框架有三层:
5.1 SchemaCache:缓存 DataTable 结构和主键列表
internal sealed class SchemaCache
{
private readonly ConcurrentDictionary<CacheKey, DataTable[]> schemaCache = new();
private readonly ConcurrentDictionary<CacheKey, List<string>> primaryKeyCache = new();
public DataTable[] GetSchema(Database database, IDbCommand command)
{
var key = new CacheKey(database.ConnectionString, command.CommandText);
if (schemaCache.TryGetValue(key, out var tables))
{
var clones = new DataTable[tables.Length];
for (int i = 0; i < tables.Length; i++)
clones[i] = tables[i].Clone();
return clones;
}
return null;
}
}
5.2 CachingMechanism + ParameterCache:缓存存储过程参数集合
调用存储过程时,DeriveParameters 会去数据库查询参数元信息,这个过程很慢。但同一个存储过程的参数是固定的,所以只需要查询一次:
internal readonly record struct CacheKey(string ConnectionString, string CommandText);
internal class CachingMechanism
{
private readonly ConcurrentDictionary<CacheKey, IDataParameter[]> paramCache = new();
public IDataParameter[] GetCachedParameterSet(Database database, IDbCommand command)
{
var key = new CacheKey(database.ConnectionString, command.CommandText);
return paramCache.TryGetValue(key, out var cached)
? CloneParameters(cached)
: null;
}
}
ParameterCache 是对 CachingMechanism 的进一步封装:
internal class ParameterCache
{
private readonly CachingMechanism cache = new();
public void SetParameters(DbCommand command, Database database)
{
var cached = cache.GetCachedParameterSet(database, command);
if (cached == null)
{
database.DiscoverParameters(command);
var copy = CreateParameterCopy(command);
cache.AddParameterSetToCache(database, command, copy);
}
else
{
foreach (var p in cached)
command.Parameters.Add(p);
}
}
}
5.3 深克隆为什么是必须的
缓存里存的是原始对象,但调用方可能会修改它们(比如 param.Value = xxx)。如果直接返回缓存引用,第二次调用的参数值就被上一次污染了。所以每次返回都必须克隆:
public static IDataParameter[] CloneParameters(IDataParameter[] originals)
{
var cloned = new IDataParameter[originals.Length];
for (int i = 0; i < originals.Length; i++)
cloned[i] = (IDataParameter)((ICloneable)originals[i]).Clone();
return cloned;
}
DataTable.Clone() 只克隆结构(列定义),IDataParameter.Clone() 克隆参数元信息——深克隆 + ConcurrentDictionary 是这个框架线程安全、永不串味的核心保证。
5.4 record struct CacheKey:多维度组合键
缓存键需要同时考虑连接字符串(不同数据库可能有同名存储过程)和命令文本(不同 SQL)。用 record struct 做组合键:
internal readonly record struct CacheKey(string ConnectionString, string CommandText);
record 自动生成了 Equals / GetHashCode,值相等的实例会被视为同一个 Key。struct 避免堆分配,readonly 保证不可变性——非常适合做 ConcurrentDictionary 的 Key。
六、DbObject 泛型基类 + IDbObject 回调
为了让实体类能方便地拿到表名、参与生命周期,框架提供了泛型基类:
public interface IDbObject
{
void OnLoaded();
}
public abstract class DbObject<T> where T : DbObject<T>, new()
{
private static readonly TableMetadata metadata = MetadataCache.GetTableMetadata<T>();
public static string DbTableName => metadata.DbTableName;
public static string BusinessTableName => metadata.BusinessTableName;
}
子类只需要继承 DbObject
var sql = $"SELECT * FROM {NeDeviceParseExecution.DbTableName} WHERE Status = @Status";
七、完整调用示例
7.1 实体定义
[DbTable("NeDeviceParseExecution")]
public sealed class NeDeviceParseExecution : DbObject<NeDeviceParseExecution>, IDbObject
{
[DbPrimaryKey]
[DbColumn("DeviceId")]
public int Id { get; set; }
[DbColumn("Status")]
public ExecutionStatus Status { get; set; }
[DbColumn("Message")]
public string Message { get; set; }
public void OnLoaded()
{
// 映射完成回调
}
}
7.2 数据查询 + 映射
var db = DbHelper.Default.Database;
var sql = $"SELECT * FROM {NeDeviceParseExecution.DbTableName} WHERE Status = @Status";
var cmd = db.GetSqlStringCommand(sql);
db.AddInParameter(cmd, "Status", DbType.Int32, (int)ExecutionStatus.Running);
using var reader = db.ExecuteReader(cmd);
var executions = new List<NeDeviceParseExecution>();
while (reader.Read())
executions.Add(ObjectMapper.MapFrom<NeDeviceParseExecution>(reader));
7.3 存储过程调用(参数缓存生效)
// 首次调用:DiscoverParameters 查数据库 -> 缓存
using var cmd1 = db.GetStoredProcCommand("sp_GetDeviceStats", 1001, "2026-01-01");
var result1 = db.ExecuteDataSet(cmd1);
// 第二次调用:命中 ParameterCache -> 直接克隆参数
using var cmd2 = db.GetStoredProcCommand("sp_GetDeviceStats", 1002, "2026-01-02");
var result2 = db.ExecuteDataSet(cmd2);
八、设计总结
这个轻量级 ORM 的核心设计理念可以归纳为三件事:
8.1 元数据只构建一次
MetadataCache.GetOrAdd 让每个 Type 的元数据(特性解析 + Expression Tree 编译)只做一次。后续所有操作都是字典查询 + 委托调用。这是 ORM 性能的基石。
8.2 所有缓存都需要克隆防线
缓存的都是原始对象(IDataParameter[]、DataTable 结构),但调用方可能会修改。框架在 SchemaCache.GetSchema、CachingMechanism.GetCachedParameterSet 两个出口统一做深克隆——返回副本,永不共享。这是线程安全的根本保证。
8.3 特性 + Expression Tree 是天生绝配
特性声明式地描述映射关系,Expression Tree 把声明翻译成高性能的强类型委托。两者组合起来,既保留了写一行特性的便利,又拥有了手写赋值的性能。
写在最后
自己造 ORM 轮子看起来像在重新发明 Dapper,但这个过程让你理解了几个之前可能模糊的问题:
- Dapper 的 SqlMapper 内部也是 Expression Tree 编译委托
- Entity Framework 的元数据缓存也是 MetadataCache 的超集
- record struct 做 ConcurrentDictionary Key 为什么比 Tuple 好
- 缓存返回克隆 vs 返回引用的线程安全边界
如果你也想在项目里嵌入一个轻量 ORM,这整个框架的源码量也就一千多行,可以直接拿来改。希望这篇文章对你有启发!