Probabilistic Databases
- 180 páginas
- 7 horas de lectura
Probabilistic databases handle uncertainty in attribute values or record presence, with applications in information extraction, RFID, scientific data management, data cleaning, data integration, and financial risk assessment. These applications generate large volumes of uncertain data best modeled by probabilistic databases. This book explores the latest in representation formalisms and query processing techniques for such data. It begins with foundational principles for representing large probabilistic databases, decomposing them into tuple-independent tables, block-independent-disjoint tables, or U-databases. The discussion then shifts to two classes of query evaluation techniques. Extensional query evaluation allows probabilistic inference to be processed within the database engine, akin to standard SQL queries, with safe queries being those that can be evaluated this way. In contrast, intensional query evaluation relies on a propositional formula known as lineage expression, applicable to all relational queries, though its data complexity can be #P-hard. The book also covers advanced topics in probabilistic data management, including top-k query processing, sequential probabilistic databases, indexing, materialized views, and Monte Carlo databases.








