Multimodal Data Systems
EvaDB is a database system for AI-powered applications that grew out of our work on exploratory video analytics and multimodal query processing. It is motivated by the observation that modern AI applications increasingly need to combine structured and unstructured data, invoke machine learning models inside data processing pipelines, and reason over outputs that are expensive to compute and difficult to optimize using traditional techniques. EvaDB studies how to provide a declarative interface for these workloads, enabling users to express end-to-end pipelines that seamlessly integrate data access, model invocation, and relational processing.
A central challenge in these settings is that machine learning operators have execution costs and selectivities that are often hard to estimate statically, while the underlying data may be high-dimensional, multimodal, and require expensive inference. EvaDB investigates how core database ideas such as physical operator selection, caching, predicate pushdown, batching, and parallel query execution can be adapted to this new setting. More broadly, it explores how a database system can serve as the execution engine for AI-native workloads, making applications over videos, images, text, and tables easier to build and scale.
Aero is a fork of EvaDB that explores adaptive query processing for machine learning workloads. It builds on EvaDB’s broader vision of AI-relational query processing, but focuses specifically on the challenge that operator costs and selectivities are often difficult to estimate statically because they depend on model behavior, intermediate inference results, and input data characteristics. Aero investigates how runtime re-optimization, feedback-driven planning, and adaptive execution strategies can make ML query processing more robust and efficient in the presence of this uncertainty.
SketchQL is a novel interface-driven system that enables users to query large video collections through visual specifications rather than low-level predicates over model outputs. It allows users to express complex video moments in terms of motion, object trajectories, and temporal structure, providing a more natural query abstraction for video retrieval.
Zeus investigates how reinforcement learning can be used to localize sparse actions in videos while avoiding the cost of exhaustive frame-by-frame inference. More broadly, Zeus studies how learning-based search policies can focus computation on the most informative regions of large video datasets, reducing inference cost while preserving retrieval quality.
Past Research Areas
Database Reliability and Debugging
SQLCheck is a tool for automatically detecting and diagnosing performance anti-patterns in SQL queries. SQL applications often encode inefficiencies such as bad logical database design and querying patterns that inhibit optimization. SQLCheck analyzes queries statically and surfaces these issues with actionable feedback, helping developers identify performance problems early rather than discovering them only after deployment.
Automated Query Reasoning and Optimization
SPES is a system for proving SQL query equivalence under bag semantics. Reasoning about whether two queries are semantically identical is a fundamental challenge in query optimization, query rewriting, and redundant computation detection. SPES brings formal reasoning techniques to this problem, enabling automated verification of query transformations that would otherwise require manual inspection or remain unchecked in practice.
Non-Volatile Memory Database Management Systems
N-Store is a lightweight DBMS for studying storage and recovery architectures for byte-addressable persistent memory. It was built as an experimental platform for evaluating NVM-aware storage engines for transaction processing workloads and for understanding how persistent memory changes the design space for logging, indexing, and data layout.
Self-Driving Database Management Systems
Peloton is a self-driving database system built to explore what it means for a DBMS to operate autonomously. It investigates how a database system can reason about workload behavior, predict future trends, and continuously adapt its configuration without manual intervention. Peloton served as a research platform for self-driving architecture, learned tuning, hybrid transactional and analytical workloads, and modern storage, indexing, and concurrency control mechanisms.
PostgreSQL-CPP is a port of the PostgreSQL Database Management System frontend to modern C++. Beyond a language translation effort, it explored how a mature DBMS codebase could be re-expressed using modern systems abstractions while preserving compatibility with existing database infrastructure. It also informed subsequent systems work and helped bootstrap parts of the Peloton system.