Designing and Implementing Data Warehouse for Agricultural Big Data
Thiết kế và triển khai kho dữ liệu cho dữ liệu lớn về nông nghiệp
Author: Vuong M. Ngo, Nhien-An Le-Khac, M-Tahar Kechadi | Year: 2019
Keywords: Business intelligent, data warehouse, constellation schema, Big Data, precision agriculture
Abstract / Summary
In recent years, precision agriculture that uses modern information and communication technologies is becoming very popular. Raw and semi-processed agricultural data are usually collected through various sources, such as: Internet of Thing (IoT), sensors, satellites, weather stations, robots, farm equipment, farmers and agribusinesses, etc. Besides, agricultural datasets are very large, complex, unstructured, heterogeneous, non-standardized, and inconsistent. Hence, the agricultural data mining is considered as Big Data application in terms of volume, variety, velocity and veracity. It is a key foundation to establishing a crop intelligence platform, which will enable resource efficient agronomy decision making and recommendations. In this paper, we designed and implemented a continental level agricultural data warehouse by combining Hive, MongoDB and Cassandra. Our data warehouse capabilities: (1) flexible schema; (2) data integration from real agricultural multi datasets; (3) data science and business intelligent support; (4) high performance; (5) high storage; (6) security; (7) governance and monitoring; (8) replication and recovery; (9) consistency, availability and partition tolerant; (10) distributed and cloud deployment. We also evaluate the performance of our data warehouse.
Bản dịch tiếng Việt:
Trong những năm gần đây, nông nghiệp chính xác sử dụng công nghệ thông tin và truyền thông hiện đại đang trở nên rất phổ biến. Dữ liệu nông nghiệp thô và sơ chế thường được thu thập thông qua nhiều nguồn khác nhau, chẳng hạn như: Internet of Thing (IoT), cảm biến, vệ tinh, trạm thời tiết, robot, thiết bị nông nghiệp, nông dân và doanh nghiệp nông nghiệp, v.v. Ngoài ra, bộ dữ liệu nông nghiệp rất lớn, phức tạp, không có cấu trúc, không đồng nhất, không chuẩn hóa và không nhất quán. Do đó, việc khai thác dữ liệu nông nghiệp được coi là ứng dụng Big Data về khối lượng, tính đa dạng, tốc độ và tính xác thực. Đây là nền tảng quan trọng để thiết lập một nền tảng trí tuệ cây trồng, cho phép đưa ra các khuyến nghị và quyết định nông học hiệu quả về tài nguyên. Trong bài báo này, chúng tôi đã thiết kế và triển khai kho dữ liệu nông nghiệp cấp lục địa bằng cách kết hợp Hive, MongoDB và Cassandra. Khả năng kho dữ liệu của chúng tôi: (1) lược đồ linh hoạt; (2) tích hợp dữ liệu từ nhiều bộ dữ liệu nông nghiệp thực tế; (3) khoa học dữ liệu và hỗ trợ kinh doanh thông minh; (4) hiệu suất cao; (5) khả năng lưu trữ cao; (6) an ninh; (7) quản trị và giám sát; (8) sao chép và phục hồi; (9) tính nhất quán, tính sẵn có và khả năng chấp nhận phân vùng; (10) phân phối và triển khai đám mây. Chúng tôi cũng đánh giá hiệu suất của kho dữ liệu của chúng tôi.
Full Document
Open in new tabDocument Ready
Click below to load the full PDF document natively.