Showing posts with label Beginner. Show all posts
Showing posts with label Beginner. Show all posts

Tuesday, December 15, 2015

Introduction to Hadoop

Dear Friends,

Couple of months back I've published a post on SQL Server 2016 new features here.

Meanwhile, let me introduce you to Hadoop. We will learn this as a series of inter-related post. So don't miss any post in between and read it serially. Let's make it fun and interesting to learn Hadoop.

So, lets understand what are the formats of data that we handle in real word.
  • Flat File
  • Rows  and Columns
  • Images and Document
  • Audio and Video
  • XML Data and many more......
Big Data is an ocean of data which an organization stores. These data come in three V's i.e. Volume, Velocity and Variety.

Now-a-days huge Volume of data are getting generated by many sources such as Facebook,Whatsapp, E-commerce sites, etc, etc, etc. These huge volume of data are getting generated with high Velocity, can say it is multiplying every seconds, every minute, every hour. Along with the huge Volume and high Velocity numerous Variety of data is generated in different forms.

These data can be in any format i.e. structure, semi-structure as well as unstructured. Data stored in the form of row and column can be well defined as structured data whereas data in form of document, image, sms, video, audio,etc can be categories into unstructured and data in html or in XML format can be semi-structure data.

Q. I am sure you must be thinking that then how does a RDBMS handles these kind of unstructured or semi-structure data in there Database?

A. Well, to handle these kind of data's we have special data type such as Varbinary(Max), XML. Drawback of this is, if we are storing an image, it is stored in binary format within the database; whereas actual image is stored in Filestream or the Server itself. Hence there is an performance impact during storing and retrieving Petabyte of data's.

Moreover to this, Big Data is not just about maintaining and growing the data year on year, but it is also about how you manages these data's to make an informative decision. Data in Big Data can also comes in various complex format, to manage and process these type of data we need large set of cluster servers.
BIG DATA & HADOOP

With this introduction to Big Data, now let me introduce you to Hadoop. 

Hadoop is a large set of cluster servers which is built to process large set of data. It has two main core component i.e. 'Hadoop MapReduce' (Processing Part) and 'Hadoop Distributed File System' (Storage Part). Hadoop project comes under Apache and that is why it is called as 'Apache Hadoop'. The idea behind these two core component came into existence when Google has released there two white paper of there project on 'MapReduce' and 'Google File System (GFS)' in the year 2004.
Hadoop was created by Doug Cutting in 2005. Cutting, He was working at Yahoo! at the time he build the software, it was named after his son's toy elephant.

Wikipedia defines Hadoop as "an open-source software framework written in Java for distributed storage and distributed processing of very large data sets on computer clusters built from commodity hardware"

Hadoop is an open source framework available in free as well as commercial use under Apache license. It allows distributing and processing of dataset across the large Cluster set. On top of this, there are lot more application build by other organisation who use Hadoop or continuously work on this product which all comes under 'Hadoop Ecosystem'. Check here to find the list of projects under Hadoop Ecosystem. As we move further we will see post on the important projects under Hadoop Ecosystem.

Apache Hadoop Architecture consist of following components:
  • Hadoop Common: It contains the libraries and other utilities needed by other module of Hadoop.
  • Hadoop Distributed File System (HDFS): It is cluster of Servers with commodity storage which is used for data storage across the cluster.
  • Hadoop YARN: This component is used for Job scheduling and Resource management in Cluster.
  • Hadoop MapReduce: The processing part of the data is done by this component. 
At least now, I can consider that you are having a fair enough idea about this technology. But how can or who is the best person to get into Hadoop??

Technical answer for that would be those who are interested to learn this technology can get in two ways:
  • As a Developer
  • As a Administrator
  1. As a Developer: Hadoop is a framework which is built in JAVA language. So having JAVA background can get easy access to become a Hadoop Developer. Since the growing popularity of Hadoop, now a days this is the most common designation you can find in Job sites. 
  2. As a Administrator: Most organisation with Hadoop installation prefer for a Part time or a
    Hadoop Administration

    Full time Administrator to manage there Hadoop Clusters. It is not compulsion that the admin should have the knowledge of JAVA to learn this technology. Indeed! they should have some basics for troubleshooting. Candidate those who are having knowledge with Database Admin (SQL Server, Oracle, etc) background who already have troubleshooting, Server maintenance, Disaster Recovery knowledge are preferred or anyone with Network or Storage or Server Admin (Windows\Linux) skills  can be the other best choice. Here in this post it is mentioned in detail who suits best for Hadoop.     
Following might be the questions in your mind if we want to get start with Hadoop Admin:

  1. Do we need any DBA skills? Of course Yes; If we need to Admin the Hadoop Cluster (Maintaining, Monitoring, Configuration, Troubleshooting,etc). 
  2. Do we need to learn Java? Yes; At least some basics to understand the Java errors while troubleshooting any issue.
  3. Do we need to understand Non-RDBMS? Yes; Hadoop understand both SQL and NoSQL (Not only SQL). So having knowledge on Non-RDBMS product is most important.
  4. Do we need to learn Linux too? Yes;  at least the basics.
In our next post we will see the concept of HDFS (Hadoop Distributed File Structure).

Interested in learning SQL Server Clustering check here. Stay tuned for many more updates...

Keep Learning and Enjoy Learning!!!

Tuesday, October 20, 2015

Error 233 - No process is on the other end of the pipe

Hi Guys,

As we have seen here an error related to SQL Server Restart. Today we will see another error i.e. Error 233 "A connection was successfully established with the server, but then an error occurred during the login process."

You might have faced this common issue in your DBA career and yes the solution is relatively simple.

So what I was doing, I was connecting to the SQL Server from Windows Login but during the login process it failed and prompted the below error message:

Error 233 in SQL Server
As I said this seems an common problem so there might be multiple workaround for this.

Now, for me the solution was to start the SQL Server management Studio (SSMS) under "Run as Administrator" (When I was fresher, I use to always wonder what difference it makes if I start any program under Run as Administrator?? Let's find out the reason for that).

Reason for this error was:

a. When we are running an application under "Run As Administrator" we get extra privileges which we may not have under the local user account.

b. So running the program under Run As Administrator will grant extra rights to the account (but it should have admin rights in Active Directory).

Once I started the SSMS under Run as Administrator and connected to SQL Server it succeeded.

There might be chances that this may not solve this  issue, check here to find more work around on the issue.

Do you know?? We should properly set the Auto Growth parameter in Database else it would end up with consuming extra space. Check here.

Thanks,
Vikas B Sahu

Keep Learning and Enjoy Learning!!!

Tuesday, August 25, 2015

Error 1814 - Could not Start SQL Server Services

Hey Guys,

Some of you might be very well familiar with the below error. Let's see under what circumstances I've faced this error.

Few months back while I was configuring "AlwaysON" on one of the Server, something went wrong and due to which I've to completely decommission the AlwaysON configuration along with Windows Clustering.

After doing this, I restated the server and when the server was up what I saw was the SQL Server service was disabled, tried starting the same through Configuration Manager but hard luck I was facing the below error message.

Error 1814
Error 1814
Let's have a look below into the SQL Server Error log what its was it looking like:

Error Log
Error Log
With these error logs we started troubleshooting:
  1. The error in the SQL Error Log states that due to insufficient disk space "TempDB" was not created.
  2. Now what we were wondering about the error, as it was a new Server and disk was having sufficient space. So insufficient space was out of question.
  3. Just to cross verify I open the My Computer tab and what I saw was only C: drive was reflecting and rest of the drive were not visible.
  4. While installation I have kept the TempDB in D: drive.
  5. Now here comes the twist, let's recollect the above activity of AlwaysON which I was performing. But for some reason I have to decommission it.
  6. During decommissioning I've put the disks into offline state and due to which the disk were not accessible.
  7. So at start of SQL Server if the TempDB is not created, the SQL Server Services will not start.
Brief note on System Databases:

  1. Master, Model, MSDB and TempDB are system databases.
  2. Files of the Master Database is used in start up parameter during the SQL Server instance starts. 
  3. So if the Master Database is corrupted SQL Server instance will not start. 
  4. Also as it's one of the important properties is it stores the location information for other Database. Therefore Master Database is said to be the heart of SQL Server.
  5. Model Database acts as a template for the other User Database as well as TempDB while creation.
  6. So even if Model Database is corrupted SQL Server instances will not start. As indirectly TempDB will not start and if TempDB will not start SQL Server instance will not start.  
So guys System Database plays an very important role in proper functioning of SQL Server Instance.

Here my question is: What if MSDB is corrupted, will the SQL Server Instance will start? You can comment below.

Thank You Guys.
Keep Learning and Enjoy Learning!!!

Thursday, July 9, 2015

Introduction to SQL Server 2016

Hello Guys,

As many of you must be aware under Microsoft flagship, they have announced SQL Server 2016 in May’15 during Microsoft Ignite conference. So, the latest available version of MS SQL Server 2016 is CTP 2.1 (Community Technology Preview) from Jun’16. Soon (Somewhere in 2016; not yet confirm) the general version RTM (Release To Manufacture) would be available.  Also the code name for this version is yet to release.


SQL Server 2016
Introducing SSMS for SQL Server 2016

Following is a brief note on the how the Software is released in phases:
  • CTP > It stands for Community Technology Preview. CTP version is released before a newly developed software is roll out in the market. It is releases to improve the software or to fix bugs present.
  • RTM > It stand for Release To Manufacture. This is the first official version which released to the customers or clients for use. 
  • CU > It stands for Cumulative Updates which keep on releasing after RTM or SP is released. Its main purpose is to fix the bug found in the product.
  • SP > It stands for Service Pack. Bunch of CU forms a SP and it is cumulative. That means say there are two SP’s are released SP1 and SP2. You can directly apply SP2 no need to apply SP1 and then SP2.
Click here to find more on release date and the version number for overall Microsoft product available from SQL Server 7.0 till SQL Server 2016.

With this information now we will see what new features are available in MS SQL Server 2016. The following diagram will give you an idea about some of the new feature in SQL Server 2016.

New Features in SQL Server 2016
New Features of SQL Server 2016

We will see in detail about these new features of SQL Server 2016 on next post here. Till then........

Keep Learning and Enjoy Learning!!!

Thank you.

Thursday, October 9, 2014

Database Restoration Error In Azure

Hi Friends,

(Started watching some more series after ending of Prison Break. GOT & Friends are among few of them.)

Today I would like to share the following error, while I was trying to restore bacpac file from one Azure Server to another Azure Server.

Requirement: There was a need to restore the bacpac file from one environment to another.

What I did was:

1 Took the backup (bacpac) file from one server.
2. While I was restoring it, it was throwing the below error:

the internal target platform type



















3. After doing some research, got to know something new, that we have to install “SQL Server Data Tools”. Below is the link to download the file:
http://msdn.microsoft.com/en-us/jj650014

4. Once it was installed, we were through the restoration process.

Thanks for reading the article. Stay Tuned.

Monday, September 8, 2014

How to create Database and Migration in SQL Azure

Hi Friends,

As we have seen the welcome note on “SQL Azure” here. This is just a continuation of that post.
Here we’ll see how to create a Server, Database as well as how to authenticate the user to access the Database in Azure.

1.       Connect to windows.azure.com.
2.       Click Subscription >> Then “Create”.
3.       Select the location (Region where you wanted to deploy your Azure Database).
4.       Create Admin Login ID & set the password.
5.       Check “Allow other windows Azure services to access this service”.
Add the list of IP’s which can able to access the server.
6.       Click “Finish”.

Once it is done. A server name will be created, which one can use to connect and create a Database.

How to create Azure Database:

1.       Once the server is configured, connect to the server. Here we can able to see only “Master” Database (Master Database is not chargeable).
2.       Under Database >> Create >> Database_Name >> Select “Edition”.
3.       Database is created.

How to migrate Other Databases to Azure:

1.       If you are planning to migrate any other RDBMS to Azure, we need to first migrate it to MS SQL 2008 then only we can migrate to Azure.
2.       Once we have migrated to generate script of Database using Azure Tool (Reason all the features are not supported if we generate the script from SQL Server).
3.       SQL Azure >> Migration Wizard >> Generate Script.
4.       Copy the script and run it to create Azure Database.
5.       Populate data using either SSIS or SQL Azure migration wizard

Max size of DB is limited to 50 GB. We can connect to windows azure portal to view the bill.

Note: Above points are my understanding from the session, there might be some flaws. Looking forward for your note to correct me if I’m wrong anywhere.   

Thanks for reading the article. Stay Tuned.


Friday, September 5, 2014

Introduction on SQL Azure

Hi Friends,

Last week I've attended a session on “SQL Azure”. So thought of sharing the key points on the blog for our all reference. Till before attending the session, it was just a word for me “AZURE” with zero knowledge on this technology (or rather my understanding was just limited to “it is something related to Cloud”).
Following are some points I’d like to share which I learned from the session:

i.    We can consider SQL Serve Azure (SSA) just like a SQL Server (SS) Version (Cloud version). 
     It is the light version of SS.
ii.   SSA is Scalable, HA, RDBMS and Secure.
iii. SSA runs only on “Windows Azure Platform”.
iv.  One can connect the server through Azure Portal as well as SSMS.
v.    SSA has two Editions. Web Edition and Business Edition with 5 GB and 50 GB DB limit     respectively.
vi.   Data can be migrated from any other RDBMS such as Oracle, MySQL, SQL Server, etc...
vii. Pricing is based on the Usage (DB Size). Large the DB size is more your billing would be.

Difference between SQL Server (SS) and SQL Server Azure (SSA):

1.  Only SQL Authentication is possible in SSA; whereas SQL as well as Windows Authentication is also possible in SS.
2.  While making login Azure does not allow login name like sa, admin, guest, administrator and root. There is no as such restriction in SS.
3.  No CPU, CAL licensing is involved in SSA like SS; rather billing is done based on consumption (More the DB Size more the bill will be).
4.  Since SSA is a light version many features are not present in SSA compared to SS.
5.  SSA table must have Cluster Index whereas it’s not compulsory in SS.
6.  Only Master DB is present in SSA on contrary Master, Model, MSDB & Tempdb is present in SS.
7.  Max size of DB in SSA is 50 GB; whereas in SS it’s in TB’s.
8.  Transaction should stay in single DB; whereas in SS it can run adhoc queries.
9.  As it is in Azure only TCP\IP protocol is supported; where as in SS it supports many protocols.
10. In Azure, Application cannot go down unexpectedly.

Guys this is just a welcome note on SQL Azure.  There are lots of things to learn in SSA. Will keep on posting as and when I‘ll have the content along with the necessary snapshots.

Note: Above points are my understanding from the session, there might be some flaws. Looking forward for your note to correct me if I’m wrong anywhere.   

Thanks for reading the article. Stay Tuned.