从零实现轻量级 ADO.NET ORM:特性驱动 + Expression Tree 委托 + 多级 ConcurrentDictionary 缓存完整架构

前言

在 .NET 生态里,ORM 选择不可谓不丰富:EF Core 功能全家桶、Dapper 极简轻量、FreeSql 国产良心……但你有没有想过——自己写一个 ORM 到底有多难?

本文就带你拆解一个生产环境真实在用的轻量级 ORM 框架,它代码量不过千行,却完整实现了:

  • 特性驱动的实体-表映射(DbTable / DbColumn / DbPrimaryKey / DbIdentity)
  • 基于 Expression Tree 动态编译的 Getter/Setter 委托(零反射开销)
  • 以 Type 到 TableMetadata 为键的元数据缓存(ConcurrentDictionary + GetOrAdd)
  • 三级缓存体系:元数据缓存 + Schema 缓存 + 参数集合缓存
  • record struct 组合键解决多维度缓存 Key 设计
  • 并发安全 + 深克隆防污染
  • IDbObject.OnLoaded() 生命周期钩子

整个架构零依赖(只依赖 System.Data.Common),可以直接嵌入任何 .NET 项目。


整体架构一览

flowchart TB
    subgraph 业务层
        A[实体类: 标记 DbTable/DbColumn]
    end

    subgraph 元数据层
        B[MetadataCache] --> C[TableMetadata]
        C --> D[ColumnMetadata[]]
        D --> E[Expression Tree Getter/Setter]
    end

    subgraph 映射层
        F[ObjectMapper] --> C
        F --> G[IDataRecord / DataRow]
        G --> F
        F --> A
    end

    subgraph 缓存层
        H[SchemaCache] --> H1[DataTable 结构缓存]
        H --> H2[主键列表缓存]
        I[CachingMechanism] --> J[ParameterCache]
        J --> K[IDataParameter 集合]
    end

    subgraph 数据访问层
        L[Database 抽象基类]
        L --> H
        L --> J
        L --> F
        M[SqlDatabase 实现] --> L
    end

    A --> B
    E --> C

五层架构从下往上:数据访问层 → 缓存层 → 元数据层 → 映射层,最上层是业务实体。每一层职责单一、完全解耦。


一、特性系统:让实体自描述

整个 ORM 的入口是一组特性(Attribute),它们告诉框架:这个类映射哪张表、这个属性对应哪个字段、哪个是主键……

[AttributeUsage(AttributeTargets.Class, Inherited = false)]
public sealed class DbTableAttribute : Attribute
{
    public string Name { get; }
    public DbObjectType ObjectType { get; }
}

[AttributeUsage(AttributeTargets.Property)]
public sealed class DbColumnAttribute : Attribute
{
    public string Name { get; }
    public bool ReadOnly { get; init; }
}

[AttributeUsage(AttributeTargets.Property)]
public sealed class DbPrimaryKeyAttribute : Attribute { }

[AttributeUsage(AttributeTargets.Property)]
public sealed class DbIdentityAttribute : Attribute { }

实际业务中的实体使用长这样:

public abstract class NeDeviceParseExecutionBase<T> : DbObject<T>
    where T : NeDeviceParseExecutionBase<T>, new()
{
    [DbPrimaryKey]
    [DbColumn("DeviceId")]
    public int Id { get; set; }

    [DbColumn("Status")]
    public ExecutionStatus Status { get; set; }

    [DbColumn("LastExecutionId")]
    public int? LastExecutionId { get; set; }

    [DbColumn("StartTime")]
    public DateTime? StartTime { get; set; }
}

[DbTable("NeDeviceParseExecution")]
public sealed class NeDeviceParseExecution : NeDeviceParseExecutionBase<NeDeviceParseExecution> { }

[DbTable("V_NeDeviceParseExecution", DbObjectType.View)]
public sealed class NeDeviceParseExecutionView : NeDeviceParseExecutionBase<NeDeviceParseExecutionView>
{
    [DbColumn("DeviceName")]
    public string Name { get; set; }
}

继承链里基类定义通用字段,子类只需要用 DbTable 指定表名——这种设计让同一张表的 Table 和 View 可以复用同一套字段映射。


二、MetadataCache:Type 到 TableMetadata 的单次编译缓存

2.1 为什么需要缓存

对一个 Type 做反射、找特性、建映射表——这些操作成本不低(几十到几百纳秒)。但 Type 的元数据在应用生命周期内是不变的,所以对每个 Type 只需要构建一次元数据,之后直接从字典里取。

2.2 ConcurrentDictionary + GetOrAdd

public static class MetadataCache
{
    private static readonly ConcurrentDictionary<Type, TableMetadata> Cache = new();

    public static TableMetadata GetTableMetadata<T>()
        => GetTableMetadata(typeof(T));

    public static TableMetadata GetTableMetadata(Type type)
    {
        return Cache.GetOrAdd(type, CreateMetadata);
    }
}

ConcurrentDictionary.GetOrAdd 完美匹配这个场景:线程安全、避免重复构建。每个 Type 的元数据只创建一次,后续全走字典查询(O(1))。

2.3 CreateMetadata:构建 TableMetadata 的完整流程

private static TableMetadata CreateMetadata(Type type)
{
    var dbTableAttribute = type.GetCustomAttribute<DbTableAttribute>();

    var columns = new List<ColumnMetadata>();
    foreach (var prop in type.GetProperties(BindingFlags.Public | BindingFlags.Instance))
    {
        var colAttr = prop.GetCustomAttribute<DbColumnAttribute>();
        if (colAttr == null) continue;

        columns.Add(new ColumnMetadata
        {
            PropertyName = prop.Name,
            ColumnName   = colAttr.Name,
            PropertyType = prop.PropertyType,
            IsNullable   = prop.PropertyType.IsNullable(),
            IsPrimaryKey = prop.IsDefined(typeof(DbPrimaryKeyAttribute)),
            IsIdentity   = prop.IsDefined(typeof(DbIdentityAttribute)),
            ReadOnly     = colAttr.ReadOnly,
            Getter       = CreateGetter(prop),
            Setter       = CreateSetter(prop)
        });
    }

    var primaryKey = columns.SingleOrDefault(c => c.IsPrimaryKey);
    var insertable = columns.Where(c => !c.ReadOnly && !c.IsIdentity).ToList();
    var updatable  = columns.Where(c => !c.ReadOnly && !c.IsPrimaryKey && !c.IsIdentity).ToList();

    var metadata = new TableMetadata
    {
        Type              = type,
        Columns           = columns,
        PrimaryKey        = primaryKey,
        InsertableColumns = insertable,
        UpdatableColumns  = updatable,
        PropertyMap       = columns.ToDictionary(c => c.PropertyName, ...),
        ColumnMap         = columns.ToDictionary(c => c.ColumnName, ...)
    };

    if (dbTableAttribute != null)
        metadata.InitializeTableInfo(dbTableAttribute.Name, dbTableAttribute.ObjectType);

    return metadata;
}

要点: - 只有标记了 DbColumn 的属性才会进入映射 - InsertableColumns / UpdatableColumns 在元数据阶段就分好组了 - PropertyMap / ColumnMap 两个字典实现了双向索引


三、Expression Tree 编译 Getter/Setter:零反射的属性访问

这是整个 ORM 最硬核的部分——用 Expression Tree 在元数据构建阶段动态编译出强类型的属性访问委托,之后的所有属性读写都是直接调用委托,完全绕开了反射。

3.1 CreateSetter:把赋值编译成委托

private static Action<object, object> CreateSetter(PropertyInfo property)
{
    var instance = Expression.Parameter(typeof(object), "obj");
    var value    = Expression.Parameter(typeof(object), "value");

    var instanceCast = Expression.Convert(instance, property.DeclaringType);
    var propertyAccess = Expression.Property(instanceCast, property);
    var valueCast      = Expression.Convert(value, property.PropertyType);
    var assign         = Expression.Assign(propertyAccess, valueCast);

    return Expression.Lambda<Action<object, object>>(
        assign, instance, value
    ).Compile();
}

对应的 IL 逻辑等价于你手写的:

Action<object, object> setter = (obj, val) => ((MyEntity)obj).MyProperty = (string)val;

CreateGetter 同理——把属性读出来再装箱成 object:

private static Func<object, object> CreateGetter(PropertyInfo property)
{
    var instance     = Expression.Parameter(typeof(object), "obj");
    var instanceCast = Expression.Convert(instance, property.DeclaringType);
    var propertyAccess = Expression.Property(instanceCast, property);
    var box = Expression.Convert(propertyAccess, typeof(object));

    return Expression.Lambda<Func<object, object>>(box, instance).Compile();
}

3.2 性能对比

方式 10 万次 Setter 说明
PropertyInfo.SetValue(反射) ~1500ms 每个属性都有装箱拆箱开销
Expression Tree 编译委托 ~12ms 快 125 倍
手写赋值代码 ~8ms 几乎追上手写

元数据只编译一次,后续全走委托调用——这是典型的一次编译、终身受益模式。


四、ObjectMapper:IDataRecord / DataRow 到实体的映射器

有了元数据缓存和委托,映射就变得非常简单:

public static class ObjectMapper
{
    public static T MapFrom<T>(IDataRecord record)
        where T : class, new()
    {
        var entity = new T();
        return SetFrom(entity, record);
    }

    public static T SetFrom<T>(T entity, IDataRecord record)
    {
        var metadata = MetadataCache.GetTableMetadata<T>();

        for (int i = 0; i < record.FieldCount; i++)
        {
            var columnName = record.GetName(i);
            var value      = record.GetValue(i);
            SetValue(entity, metadata, columnName, value);
        }

        if (entity is IDbObject callback)
            callback.OnLoaded();

        return entity;
    }

    private static void SetValue<T>(T entity, TableMetadata metadata,
                                    string columnName, object value)
    {
        if (!metadata.ColumnMap.TryGetValue(columnName, out var column))
            return;

        if (value is null || value == DBNull.Value)
        {
            if (!column.IsNullable) throw new InvalidOperationException(...);
            column.Setter(entity, null);
            return;
        }

        if (!column.PropertyType.IsInstanceOfType(value))
            value = value.ConvertTo(column.PropertyType);

        column.Setter(entity, value);
    }
}

设计亮点:

  1. 双入口:IDataRecord(DbDataReader 流式读取)和 DataRow(DataSet 离线数据)都能映射
  2. ColumnMap 索引:SELECT 里没出现的字段自动跳过
  3. IDbObject.OnLoaded():映射完成后回调

调用方一行代码搞定:

var reader = database.ExecuteReader("SELECT * FROM NeDeviceParseExecution");
var list = new List<NeDeviceParseExecution>();
while (reader.Read())
    list.Add(ObjectMapper.MapFrom<NeDeviceParseExecution>(reader));

五、多级缓存体系:不止是元数据

生产环境的 ORM 缓存应该是多层次的。我们的框架有三层:

5.1 SchemaCache:缓存 DataTable 结构和主键列表

internal sealed class SchemaCache
{
    private readonly ConcurrentDictionary<CacheKey, DataTable[]> schemaCache = new();
    private readonly ConcurrentDictionary<CacheKey, List<string>> primaryKeyCache = new();

    public DataTable[] GetSchema(Database database, IDbCommand command)
    {
        var key = new CacheKey(database.ConnectionString, command.CommandText);
        if (schemaCache.TryGetValue(key, out var tables))
        {
            var clones = new DataTable[tables.Length];
            for (int i = 0; i < tables.Length; i++)
                clones[i] = tables[i].Clone();
            return clones;
        }
        return null;
    }
}

5.2 CachingMechanism + ParameterCache:缓存存储过程参数集合

调用存储过程时,DeriveParameters 会去数据库查询参数元信息,这个过程很慢。但同一个存储过程的参数是固定的,所以只需要查询一次:

internal readonly record struct CacheKey(string ConnectionString, string CommandText);

internal class CachingMechanism
{
    private readonly ConcurrentDictionary<CacheKey, IDataParameter[]> paramCache = new();

    public IDataParameter[] GetCachedParameterSet(Database database, IDbCommand command)
    {
        var key = new CacheKey(database.ConnectionString, command.CommandText);
        return paramCache.TryGetValue(key, out var cached)
            ? CloneParameters(cached)
            : null;
    }
}

ParameterCache 是对 CachingMechanism 的进一步封装:

internal class ParameterCache
{
    private readonly CachingMechanism cache = new();

    public void SetParameters(DbCommand command, Database database)
    {
        var cached = cache.GetCachedParameterSet(database, command);
        if (cached == null)
        {
            database.DiscoverParameters(command);
            var copy = CreateParameterCopy(command);
            cache.AddParameterSetToCache(database, command, copy);
        }
        else
        {
            foreach (var p in cached)
                command.Parameters.Add(p);
        }
    }
}

5.3 深克隆为什么是必须的

缓存里存的是原始对象,但调用方可能会修改它们(比如 param.Value = xxx)。如果直接返回缓存引用,第二次调用的参数值就被上一次污染了。所以每次返回都必须克隆:

public static IDataParameter[] CloneParameters(IDataParameter[] originals)
{
    var cloned = new IDataParameter[originals.Length];
    for (int i = 0; i < originals.Length; i++)
        cloned[i] = (IDataParameter)((ICloneable)originals[i]).Clone();
    return cloned;
}

DataTable.Clone() 只克隆结构(列定义),IDataParameter.Clone() 克隆参数元信息——深克隆 + ConcurrentDictionary 是这个框架线程安全、永不串味的核心保证。

5.4 record struct CacheKey:多维度组合键

缓存键需要同时考虑连接字符串(不同数据库可能有同名存储过程)和命令文本(不同 SQL)。用 record struct 做组合键:

internal readonly record struct CacheKey(string ConnectionString, string CommandText);

record 自动生成了 Equals / GetHashCode,值相等的实例会被视为同一个 Key。struct 避免堆分配,readonly 保证不可变性——非常适合做 ConcurrentDictionary 的 Key。


六、DbObject 泛型基类 + IDbObject 回调

为了让实体类能方便地拿到表名、参与生命周期,框架提供了泛型基类:

public interface IDbObject
{
    void OnLoaded();
}

public abstract class DbObject<T> where T : DbObject<T>, new()
{
    private static readonly TableMetadata metadata = MetadataCache.GetTableMetadata<T>();

    public static string DbTableName => metadata.DbTableName;
    public static string BusinessTableName => metadata.BusinessTableName;
}

子类只需要继承 DbObject,就自动拥有了 DbTableName 静态属性:

var sql = $"SELECT * FROM {NeDeviceParseExecution.DbTableName} WHERE Status = @Status";

七、完整调用示例

7.1 实体定义

[DbTable("NeDeviceParseExecution")]
public sealed class NeDeviceParseExecution : DbObject<NeDeviceParseExecution>, IDbObject
{
    [DbPrimaryKey]
    [DbColumn("DeviceId")]
    public int Id { get; set; }

    [DbColumn("Status")]
    public ExecutionStatus Status { get; set; }

    [DbColumn("Message")]
    public string Message { get; set; }

    public void OnLoaded()
    {
        // 映射完成回调
    }
}

7.2 数据查询 + 映射

var db = DbHelper.Default.Database;

var sql = $"SELECT * FROM {NeDeviceParseExecution.DbTableName} WHERE Status = @Status";
var cmd = db.GetSqlStringCommand(sql);
db.AddInParameter(cmd, "Status", DbType.Int32, (int)ExecutionStatus.Running);

using var reader = db.ExecuteReader(cmd);
var executions = new List<NeDeviceParseExecution>();
while (reader.Read())
    executions.Add(ObjectMapper.MapFrom<NeDeviceParseExecution>(reader));

7.3 存储过程调用(参数缓存生效)

// 首次调用:DiscoverParameters 查数据库 -> 缓存
using var cmd1 = db.GetStoredProcCommand("sp_GetDeviceStats", 1001, "2026-01-01");
var result1 = db.ExecuteDataSet(cmd1);

// 第二次调用:命中 ParameterCache -> 直接克隆参数
using var cmd2 = db.GetStoredProcCommand("sp_GetDeviceStats", 1002, "2026-01-02");
var result2 = db.ExecuteDataSet(cmd2);

八、设计总结

这个轻量级 ORM 的核心设计理念可以归纳为三件事:

8.1 元数据只构建一次

MetadataCache.GetOrAdd 让每个 Type 的元数据(特性解析 + Expression Tree 编译)只做一次。后续所有操作都是字典查询 + 委托调用。这是 ORM 性能的基石。

8.2 所有缓存都需要克隆防线

缓存的都是原始对象(IDataParameter[]、DataTable 结构),但调用方可能会修改。框架在 SchemaCache.GetSchema、CachingMechanism.GetCachedParameterSet 两个出口统一做深克隆——返回副本,永不共享。这是线程安全的根本保证。

8.3 特性 + Expression Tree 是天生绝配

特性声明式地描述映射关系,Expression Tree 把声明翻译成高性能的强类型委托。两者组合起来,既保留了写一行特性的便利,又拥有了手写赋值的性能。


写在最后

自己造 ORM 轮子看起来像在重新发明 Dapper,但这个过程让你理解了几个之前可能模糊的问题:

  • Dapper 的 SqlMapper 内部也是 Expression Tree 编译委托
  • Entity Framework 的元数据缓存也是 MetadataCache 的超集
  • record struct 做 ConcurrentDictionary Key 为什么比 Tuple 好
  • 缓存返回克隆 vs 返回引用的线程安全边界

如果你也想在项目里嵌入一个轻量 ORM,这整个框架的源码量也就一千多行,可以直接拿来改。希望这篇文章对你有启发!