QCATS:查询感知Transformer切片高效预测查询处理
Get the story
2026年10月8日,arXiv Databases(一手)发表QCATS框架,提出面向数据库系统内高效稀疏推理的查询上下文感知Transformer切片方案。该框架在查询粒度执行,利用查询谓词与元数据统计,在模型执行前预选上下文对齐的FFN切片,并采用异步CPU-GPU流水线与路由感知批处理优化。实验在BERT-base与Qwen-0.6B上的四个预测查询负载进行,结果显示延迟最高降低4.42倍,预测精度与稠密基线相当。目前未见与其他报道矛盾之处。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · DatabasesQCATS:面向高效预测查询处理的查询上下文感知 Transformer 切片
QCATS 是一种查询上下文感知 Transformer 切片框架,在数据库系统内实现高效稀疏推理。它在查询粒度执行,用查询谓词和元数据统计在模型执行前预选上下文对齐的 FFN 切片,并采用异步 CPU-GPU 流水线与路由感知批处理优化。在 BERT-base 和 Qwen-0.6B 的四个预测查询负载上,延迟最高降低 4.42x,预测精度与稠密基线相当。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.