B.TechSemester 62023-24Big DataKCS061

Big Data (KCS061) - AKTU Question Paper 2023-24

B.Tech · Semester 6 · Free PDF Download

This is the official AKTU Big Data Previous Year Question Paper for B.Tech Semester 6, academic session 2023-24. Published by Dr. A.P.J. Abdul Kalam Technical University (AKTU/UPTU), Lucknow. Free PDF download — no login required.

Course:B.Tech
Semester:Semester 6
Session:2023-24
University:AKTU / UPTU

Rate this paper

Questions Asked in 2023-24

Big Data (KCS061) — complete question paper · 100 marks · 3 Hours

Section AAttempt all q u e s t i o n s i n b r i e f . 2 x 10 = 20
  • a
    What are the different types of digital data commonly encounter ed in Big Data applications? Provide examples of structured, semi-structured, and unstructured data
  • b
    What constitutes a Big Data platform? 02 1
  • c
    What is Hadoop Streaming? 02 2
  • d
    Discuss the data formats commo nly used in Hadoop environments. 02 2
  • e
    Describe the concepts of file sizes, block sizes, and block ab straction in HDFS
  • f
    What are the benefits and challenges of using HDFS for distrib uted storage and processing?
  • g
    What are the characteristics and use cases for schedulers such as Fair Scheduler and Capacity Scheduler?
  • h
    What is YARN? 02 4 i. What is Apache Pig? 02 5 j. Describe the Grunt shell in Apache Pig. 02 5
Section BAttempt any three o f t h e f o l l o w i n g : 3 x 10 = 30
  • a
    Distinguish between data anal ysis and reporting in the conte xt of Big Data. How does advanced analytics go beyond traditional reporti ng to uncover hidden patterns, trends, and correlations in data?
  • b
    Explain Apache Hadoop and its role in big data processing. W hat are the core components of the Apache Hadoop ecosystem, and how do they work together to enable distributed data storage and processing?
  • c
    Explain the core concepts of HDFS, including NameNode, DataN ode, and the file system namespace. How do these components work together to manage data storage and replication in Hadoop clusters?
  • d
    Define NoSQL databases. What are the key characteristics and benefits of NoSQL databases compared to traditional relational databases?
  • e
    Provide an overview of Apache Hive architecture and its comp onents. How does Hive translate SQL-like queries into MapReduce jobs for data processing in Hadoop?
Section CAttempt any one p a r t o f t h e f o l l o w i n g : 1 x 10 = 10
  • a
    Describe the "5 Vs" of Big Data. What do each of these terms represent in the context of Big Data, and why are they essential considerations for data management and analysis?
  • b
    Provide examples of real-world applications where Big Data a nalytics have been instrumental. How do industries such as healthcare, finance, e-commerce, and transportation leverage Big Data to gain insigh ts and create value?
  • a
    Describe the Hadoop Distribute d File System (HDFS). How does HDFS manage the storage and replication of data across a distributed cluster of machines?
  • b
    Discuss the process of devel oping a MapReduce application. W hat are the key steps involved in writi ng, testing, and deploying a Map Reduce program?
  • a
    Explain how HDFS stores, reads, and writes files. Describe t he sequence of operations involved in storing a file in HDFS, retrieving da ta from HDFS, and writing data to HDFS
  • b
    Describe the considerations f or deploying Hadoop in a cloud environment. What are the advantages and challenges of running Hadoop clusters on cloud platforms like Amazon Web Services (AWS), Microsoft Azure, and Goo gle Cloud Platform (GCP)?
  • a
    Explain the operations for creating, updating, and deleting documents in MongoDB. What are the MongoDB CRUD operations, and how are y used to manipulate data in collections?
  • b
    Discuss Resilient Distributed Datasets (RDDs) in Spark. What a r e RDDs, and how do they enable fault-tolerant and distributed dat a processing in Spark applications?
  • a
    Introduce the concepts of HBase and its role in the Hadoop e cosystem. How does HBase differ from traditi onal relational databases, and what advantages does it offer for storing and accessing large-scale data?
  • b
    Discuss the HiveQL language used in Apache Hive. How does Hi veQL support SQL-like syntax for defining tables, querying data, and performing data manipulation operations?

Question text is extracted from the official AKTU question paper PDF above. Hindi translations are omitted — every question is printed in English in the original paper. Last verified: 2026-08-23.

Big Data — Other Year Papers

AKTU Big Data PYQs from other sessions