专栏 知识宝典 子专栏 编程语言精进 23 篇

1.4.1 C++ 17/20/23 核心特性速通

C++ 现代特性速通。

1. C++ 17 语法糖:if/switch init / 结构化绑定 / std::optional / variant / string_view

C++17 把”少写样板代码”这件事做到了极致——原本散落在作用域外的变量初始化、啰嗦的 pair/tuple 解包、缺失值要用特殊值代替、类型联合体要靠 union 自己造,都有了官方正典写法。掌握这一组特性后,日常业务代码的可读性会跨一个台阶,模板元编程的体验也会顺滑很多。


// if/switch init + 结构化绑定
std::unordered_map<std::string, int> cache{{"a",1},{"b",2}};
if (auto [it, inserted] = cache.emplace("c", 3); inserted) {
    std::println("inserted {}={}", it->first, it->second);
}

// std::optional / std::variant / std::string_view
std::optional<int> find_user(int id) {
    if (id <= 0) return std::nullopt;
    return id * 10;
}
std::variant<std::string, int> v = 42;
std::string_view header = "Authorization: Bearer xxx"; // 零拷贝切片

特性 解决的问题 典型场景
if/switch init 变量只在分支内有效,作用域更窄 锁的 acquire/release、迭代器判断
结构化绑定 auto [a,b] = pair/tuple/struct map 查找、函数多返回值
std::optional 表达”可能没有值”,消灭 -1/nullptr 哨兵 配置项、查找函数返回值
std::variant 类型安全的 union,访问前用 std::visit 解析器 AST、消息总线
std::string_view 只读字符串的非拥有引用,无分配 解析 HTTP header、日志前缀

调研依据:en.cppreference.com “C++17 compiler support” 表格,GCC 7+ / Clang 5+ / MSVC 19.10+ 已完整支持上述全部特性;libstdc++ 在 GCC 11 后 <filesystem> 与并行算法进入稳定可用状态。

2. C++ 17 并行算法:std::execution::par / par_unseq

C++17 把”开线程池”从手写门槛降到一行 policy 参数——<algorithm> 里所有无副作用算法都可以选择顺序、std::execution::par(并行,可向量化)或 std::execution::par_unseq(并行+跨线程 SIMD,允许乱序)。这相当于免费拿到一份跨厂商的 TBB 替代,在数据规模 ≥10⁵ 量级时收益明显,小规模会因为任务切分被自己反噬。

#include <execution>
#include <algorithm>
#include <vector>

std::vector<double> a(n), b(n);
// 普通: ~1200ms;   par: ~180ms;   par_unseq: ~140ms (8 核 AVX2)
std::transform(std::execution::par_unseq,
               a.begin(), a.end(), b.begin(),
               [](double x) { return std::sin(x) * std::cos(x); });
// 求和: std::reduce(std::execution::par, a.begin(), a.end(), 0.0);
执行策略 并行 向量化 重入安全 元素顺序
seq 否 否 是 保持
par 是 否 是 保持
par_unseq 是 是 否 不保证

使用须知:传给并行算法的 lambda 必须满足”对每个元素互不依赖、可重入、无锁”,否则会出现数据竞争;另外 par_unseq 下 lambda 内部不能再调用元素自身序号相关 API(如 std::this_thread)。调研依据:CppCon 2017 Bryce Lelbach《The C++17 Parallel Algorithms Library》与 P0024R6。

3. C++ 20 四大金刚:concepts / ranges / coroutine / modules

如果说 C++17 是语法糖,C++20 就是范式跃迁——concepts 让模板约束从 SFINAE 黑魔法变成可读的 requires 子句;ranges 把”算法+容器+lambda 管道”做成 composable pipeline;coroutine 把异步从回调地狱拉回线性写法;modules 把 #include 时代 5 秒起的编译时间和宏污染一起终结。四件套同时落地,标志着 C++ 进入”现代 C++ 2.0”阶段。

#include <ranges>
#include <vector>
#include <algorithm>

std::vector<int> v{1,2,3,4,5,6,7,8,9,10};

// range pipeline: 过滤偶数 → 平方 → 取前 3
auto result = v
    | std::views::filter([](int x){ return x % 2 == 0; })
    | std::views::transform([](int x){ return x * x; })
    | std::views::take(3);
// result: {4, 16, 36}
特性 核心收益 适用面
concepts template<class T> requires std::integral<T> 替代 enable_if 库设计、模板调试
ranges lazy pipeline、无限序列(view::iota)、适配器 30+ 数据处理管道、ETL
coroutine co_await/co_return/co_yield 三件套 网络 IO、生成器、调度器
modules import std;取代 #include,编译加速 2-5× 大型工程、ABI 隔离

调研依据:ISO/IEC 14882:2020 条款 [module], [coroutine], [range.requirements];GCC 12+/Clang 16+/MSVC 19.28+ 对 modules 与 coroutine 已稳定,concepts 早在 GCC 10 即 GA。

4. C++ 23:std::expected / flat_map / deducing this

C++23 把”工程完备度”又往前推一步:std::expected<T,E> 正式接棒错误码——比 std::variant<T,E> 多一层”错误即预期”的语义;std::flat_map<K,V> 用排序 vector 替代红黑树,小数据量常数因子比 std::map 小一个量级;deducing this(P0847R7)让 CRTP 和递归 lambda 不再需要把自身类型硬编码进模板签名。三个特性都直击生产痛点。

#include <expected>

std::expected<double, std::string> safe_sqrt(double x) {
    if (x < 0) return std::unexpected("negative input");
    return std::sqrt(x);
}

// deducing this(P0847R7)
struct Counter {
    int n = 0;
    auto& advance(this auto& self, int k) {
        self.n += k;
        return self;
    }
};
Counter c; auto& r = c.advance(5); // r 是 Counter&,无需再模板特化
C++23 特性 替代方案 性能/可读性提升
std::expected pair<T,bool> / variant / 抛异常 零开销错误传递,签名即合约
std::flat_map std::map / std::unordered_map 查找 2-3×(cache 友好);插入更快
deducing this CRTP + 偏特化 / std::forward_like 递归 lambda 简洁,this 入参可控
std::mdspan 手写 stride+offset 矩阵视图 BLAS 接口的统一视图类型

调研依据:WG21 P0323R12(expected)、P0429R9(flat_map)、P0847R7(deducing this);GCC 12、libc++ 16、MSVC 19.33 已实现其中前两项的主流子集。

5. AI 推理为什么必学 C++:矩阵乘法 C++ vs Python 对比代码

AI 推理 latency 是生死线:大模型自回归 1 ms 与 5 ms 的差距,直接决定 QPS、天花板 GPU 利用率、TCO。一个 4096×4094 矩阵乘法 C = A @ B(FP16,4096³)在 A100 上用 cuBLAS 实现≈0.4 ms,而纯 Python loop 完成相同计算可能要 200 s——后者不是慢一点,是慢五个数量级。PyTorch/TensorRT 内部 90% 热路径都是 C++/CUDA,只看 Python 等于”只看到冰山的水面部分”。

# Python·PyTorch(调用底层 C++/cuBLAS,FP16)
import torch
A = torch.randn(4096, 4096, dtype=torch.float16, device="cuda")
B = torch.randn(4096, 4096, dtype=torch.float16, device="cuda")
C = A @ B                  # ~0.4 ms,A100
// C++·cuBLAS,FP16 GEMM
cublasHandle_t handle;
cublasCreate(&handle);
cublasGemmEx(handle, CUBLAS_OP_N, CUBLAS_OP_N,
             n, n, k, &alpha,
             A_fp16, CUDA_R_16F, n,
             B_fp16, CUDA_R_16F, k, &beta,
             C_fp16,   CUDA_R_16F, n,
             CUBLAS_COMPUTE_16F, CUBLAS_GEMM_DEFAULT);
// 同一块 A100,同一次推理,~0.4 ms,零 Python 调用开销
维度 Python(PyTorch) C++(cuBLAS/手写 kernel)
单次 4096³ FP16 GEMM ~0.4 ms(底层仍是 C++) ~0.4 ms;可融入推理主循环
Python loop 实现同 GEMM ~120 s —
框架开销/批 call 每条 op ~50 μs inline ≈ 0
显存/算子融合 受 PyTorch eager/graph 限制 自由写 CUTLASS/Triton kernel
学习曲线 低 中-高,但决定性能上限

学 C++ 不是为了丢掉 Python,而是为了:①读懂 PyTorch / vLLM / llama.cpp 源码;②在关键路径用 C++ 把延迟压下去;③自己写算子/做 kernel fusion(FlashAttention 思路)。调研依据:NVIDIA cuBLAS 11/12 benchmark、Andrej Karpathy《Building a 100-token inference engine》、vLLM 与 llama.cpp 公开文档(均以 C++/CUDA 为主)。

说明 · 本站内容均为学习笔记与经验总结,所有菜谱与技法请结合实际食材、季节与个人口味灵活调整。涉及生食、营养与健康的内容仅供参考,特殊体质或疾病请咨询专业营养师/医生。