Posts

Showing posts with the label Distributed System

[Distributed File System] MapReduce: simplified data processing on large clusters Overview

This is a paper summary of the paper, "MapReduce: simplified data processing on large clusters". What is the paper trying to do? Inspired by the map and reduce primitives in functional languages, the paper introduces a new abstraction whereby map and reduce operations allows the program to parallelize large computations easily. In short, it allows the anyone to execute programs with parallelization, fault-tolerance, data distribution and load balancing without bothering with the mess that usually comes with it. What do you think is the contribution of the paper? The major contribution of this paper is developing an interface that distributes large-scale computations using MapReduce. This allows it to achieve ‘automatic parallelization’. Another big contribution is implementing this interface on large clusters of commodify PCs. What are its major strengths? Model is easy-to-use. Because it hides away the details of parallelization, fault-tolerance, locality optimiza...

[Distributed File System] A Berkeley View on Serverless Computing Overview

This blog is a paper summary of the paper, "A Berkeley View on Serverless Computing". What is the paper trying to do? The paper serves as a good introduction to serverless computing. It gives an introduction to serverless computing then goes on to talk about how it started (the motivations), the limitations and what the authors predict serverless computing will become in the future. It does a good job in explaining how the serverless cloud handles virtually all the system administration operations and makes it easier for programmers do what they usually do on the cloud. What do you think is the contribution of the paper? Again, the paper’s major contribution is its very detailed introduction of server less computing. What are its major strengths? The major strengths of serverless computing are: The appearance of infinite computing resources on demand. The elimination of an up-front commitment by cloud users. The ability to pay for use of computing resources on...

[Distributed File System] Dynamo: Amazon's Highly Available Key-value Store Overview

This is a paper summary of the paper, "Dynamo: Amazon's Highly Available Key-value Store". What is the paper trying to do? This paper is trying to “present the design and implementation of Dynamo, a highly available key-value storage system”, that is used in Amazon’s core services. By sacrificing consistency under certain failure scenarios, Dynamo is able to reach an “always-on” experience in terms of availability. This allows it to be successful in handling server failures, data center failures and network partitions. Additionally, Dynamo is incrementally scalable and can scale up and down while it is up and running. Dynamo uses a combination of technologies, including extensive use of object versioning, application-assisted conflict resolution, data partitioning, replication via consistent hashing. Additionally, during updates, quorum-like technique and a decentralized replica synchronization protocol is used to maintain consistency amongst replicas. What do yo...

[Distributed File System] Introduction to Ceph

Image
Ceph 1. Preface This blog focuses on the paper, Ceph: A Scalable High-Performance Distributed File System. Note that this paper was written in 2006 and the implementation of it might be different from what is described in the paper (and here). Details on more in-depth concepts, such as CRUSH, Metadata Server and Object Storage Device cluster are skipped in this blog (might include it in future blogs). 2. Introduction Ceph is distributed file system that builds on several philosophical and design principles. 2.1. Philosophical Principles Starting with philosophical principles, Ceph is open source . This means that it is free to use, free from being vendor specific, free to modify and free to share. Secondly, Ceph is Community-focused means that anybody can decide future step, anybody can fix a bug and update the documentation because all of us as a whole are smarter than some of us, so we can end up with better product. 2.2. Design Principles First, Ceph is Scalable . ...

[Cloud] OpenShift vs Kubernetes

Image
OpenShift VS Kubernetes Openshift is based on Kubernetes and docker. In other words, OpenShift is a modded version of Kubernetes. Below is an example of the namespace component of Kubernetes. As you can see, OpenShift replaces some of the original Kubernetes components with their own. Below is a more extensive list of differences: https://www.whizlabs.com/blog/wp-content/uploads/2019/08/openshift-vs-kubernetes-table.png

[Cloud] OpenShift vs OpenStack

1. OpenShift vs OpenStack OpenStack turns servers into cloud . It can be used to automate resource allocation so customers can provision virtual resources. OpenShift is a container centric model that leverages core concepts of Kubernetes and packages them in a neat way for developers to deploy applications on the cloud. 1.1. Concerning Containers OpenStack typically uses hypervisors like KVM, Xen or VMware to spin up virtual machines. On the other hand, OpenShift can run bare metal or it may run on Virtual Machines but it always uses containers on top of them. The containerization technology that they use is almost exclusively Docker. (Note: OpenStack does offer containerization support as well, it is meant to be used more of less like VPS and is optional.) 1.2. Distributed System OpenStack is not exclusively a distributed system . It can take control over an entire data center but that’s nowhere as global as a Kubernetes cluster. You would need a lot of e...

[Cloud] OpenStack, Magnum, OpenShift

Image
1. OpenStack OpenStack an open-source cloud operating system that turns your server into cloud environments. In other words, it provides an open alternative to the top cloud providers. It is IaaS, and it can be used to automate resource allocation so customers can provision virtual resources like VPS, block storage, object storage among other things.   2. Magnum Magnum is an OpenStack API service that makes container orchestration engines , such as Docker Swarm, Kubernetes, and Mesos available, a first class resources in OpenStack. Magnum uses Heat to orchestrate an OS image, which contains Docker and Kubernetes and runs that image in either virtual machines or bare metal in a cluster configuration. 3. OpenShift OpenShift is a platform as a service (PaaS) that leverages the core concepts of Kubernetes and packages them in a neat way for developers to deploy applications on the cloud. In short, it’s a modded Kubernetes, and accepts kubctl commands. ...

[Cloud] Amazon EKS Overview

Image
1. Amazon EKS 1.1. Overview Amazon EKS is a managed service that helps make it very easy to run Kubernetes on AWS . The idea is that most applications will run on EKS with minimal mods, if any.   Through EKS, organizations can run Kubernetes without cumbersome steps, such as: Creating the Kubernetes master cluster Configuring service discovery, Kubernetes primitives Porting and Creating database instances Setting up load balancing (eg. with HA proxy) Security Networking Hosting Control Planes across different availability zones to prevent single point of failure (Highly Available) Managing Control Plane, so users do not need to worry about components like etcd, kube-controller-manager, kube-apiserver, cloud-controller-manager and kube-scheduler.   Basically EKS = Kubernetes-as-a-service   1.2. Running Kubernetes without EKS: Manual deployment on EC2 IT teams can run a self-hosted Kubernetes environment on an EC2 instance. Deploy wi...

[Cloud] Kubernetes Overview

Image
1. Kubernetes 1.1. Overview Kubernetes is an open-source system that allows organizations to deploy and manage containerized applications like platforms as a service (PaaS), batch processing workers, and microservices in the cloud at scale. Through an abstraction layer created on top of a group of hosts, development teams can let Kubernetes manage a host of functions--including load balancing, monitoring and controlling resource consumption by team or application, limiting resource consumption and leveraging additional resources from new hosts added to a cluster, and other workflows. 1.2. Kubernetes Architecture (Master, Worker)   The Kubernetes master is responsible for maintaining the desired state for your cluster . The master can also be replicated for availability and redundancy . When you interact with Kubernetes, eg. via the kubectl command-line interface, you’re communicating with the master. The worker nodes in a cluster are the machines (VMs, ph...