Sunday, September 22, 2019

Starting a SRE (Site Reliability Engineering) Practice

SRE & Your Business 

When considering implementing an SRE department, it's one thing to read Google's SRE book and implement it the way they did in their IT centric organization.   I've found the implementation to be a bit different in organizations where IT is not the entire driver behind the business. I think it's very likely to be more of a challenge to implement in an environment where IT could be considered more of a utility to the business rather than the business moneymaker. Frankly, the majority of businesses out there are like this. In my experience, I discovered that a SRE team can encounter a fair bit of push-back to new 'priorities', 'automation ideas', and 'procedures' from other IT departments within the enterprise that have already established IT operation procedures and routines. Our team needed to evangelize and prove themselves as an SRE department before getting traction in getting our agenda considered with other teams.

Monitoring 

Members of new SRE teams might be excited about all the opportunities for automating in an organization. As a fledgling SRE team, you actually might find that much of your early automation ends up revolving around the monitoring and triage of production incidents. After all, the 'R' in SRE stands for Reliability. One of the first things our team did that created value beyond our practice in the enterprise was create monitoring dashboards from tools like AppDynamics and Splunk. These dashboards displayed and alerted on pertinent, inter-team SLO's and SLA's - many of which we derived from production incidents we tracked. Tracking inter-team SLO's was important because we discovered that inter-team incidents/problems fell through the cracks until our SRE team started owning them. Correlating each production incident to a monitor/alert we created helped us ensure: 

  • We could be proactive if similar circumstances aligned towards another analogous incident occurring, and 
  • We could likely pinpoint how to resolve the new issue because it had happened in the past and had an existing incident ticket and past resolution. 
Keeping an automated eye on the enterprise systems in production, we discovered more insights intrinsic to reliability in the organization:
  • We began to anticipate and predict new production incidents. This gave us a leg up on being proactive in preventing the incident, or at the very least, the resolution. Managers and operations teams found this information valuable which gave us a foot in the door for further discussions to socialize our agenda. 
  • We began to see correlating alerts across disparate departments associated with specific incidents, which allowed us to drive the problem to a root cause, and in some cases, several root causes with contributing factors in disparate configuration, architecture, and code. 
Automation from an operations/monitoring perspective definitely trumped any automation we did to improve IT deployments.

Challenges to Reliability 

In an organization where the business is not IT centric, Google's concept of the Error Budget sounded great, but was much more difficult to enforce, or even get any kind of commitment on. The main reason was really quite simple - the business that made the money drove what features/projects went into production and when, no holds barred. The business held the trump card - they made the money - and so long as we could keep things running that paradigm didn't seem like it was going to change. Other challenges to reliability that we ran into early on included:

  • Either non-existent or out of date documentation which caused us problems on multiple fronts, including network upgrades and developing/consuming internal APIs for application services. 
  • Configuration management issues Issues with data integrity across multiple systems. 
  • Lack of discipline and followup surrounding post mordems 

Collaboration 

Once high priority incidents were discovered that weren't owned by teams or a department, our team took ownership of those incidents and drove them through to resolution. This generally meant communicating and working towards the resolution with many different teams across the enterprise, and helping those teams understand the root cause and enabling/empowering those teams to build the fix & resolution. This kind of cooperation allowed us to get a good reputation within the organization and helped the teams responsible for the root cause feel like they also 'owned' and contributed to the solution in a positive, empowering way.

Expectations 

Getting positive results (ROI) from a new SRE team doesn't happen overnight. Unless you've handpicked seasoned developers from your existing development/operations team, it will take time for your newbies (even if they have past experience as an SRE) to get a grasp on your business & technology domain. Their ramp-up time may depend on how you've structured the implementation of the team itself. Some questions that might have a bearing on this:

  • Is it an entity on its own, or embedded across the development/operations teams? Google and Facebook use embedded team models. We started with an autonomous team, but began to head towards an embedded model. 
  • Are other teams aware of what the SRE team will be doing and how this new team can potentially help them? 
  • How current is your documentation?

Tuesday, January 8, 2019

The Job the Wasn't There - A Lesson in Applying for Jobs

Hermione (made up name to protect her identity) was a student in my fast-track Web Developer class at SAIT in early 2017.  Like a lot of students I get in that course, she was anxious about obtaining work after receiving her diploma.

One day in between exercises, I showed the class several 'Careers' web pages of good, local web design companies.  One of those companies was Critical Mass, a company I had actually consulted with before.  I often recommend this company to students because they have a world class client list, they do internships, and I have experience with them.  That particular day, they happened to have an opening at the time for a Junior Web Designer, but no posted opportunities for internships.  I encouraged the students to apply for the Junior Web Designer opportunity and Hermione challenged me...

"How can we do that when we don't have all the qualifications in their list of requirements?"

I often get this question, and I had an answer. "You need to understand how a company creates a job description.  Many put it together as a list of qualifications for the perfect candidate.  Others will build the job description based on an existing successful employee in the company. They realize that most of the applicants won't match all of the qualifications - and this is particularly true in the IT industry. "

Hermione digested my answer, and piped up again. "But we're still in school and we have several more weeks before we'd be available to start working!  Does it really make sense to apply now for a position like this?"

"Absolutely!" I replied. "You never know what might come out of an application.  The hiring process for many companies takes several weeks.  There's usually a bunch of interviews for them to schedule and have, and then some planning and logistics around actually bringing the successful applicant aboard.  You never know what will happen out of an application."

She still looked skeptical.  I moved the class onto another exercise and didn't think too much more about it.

Several weeks later, I received the following email from Hermione:

"I took your advice about applying for jobs and I applied at Critical Mass for a 
Typing letters back and forth about a job opportunity
Photo by Kaitlyn Baker on Unsplash
Junior Web Developer position knowing that I was NOT qualified and that they probably never call me back. Guess what? They called me back! They don't think I am ready for the Junior Web Developer position, but they want me to interview for their internship program. The interview is on . Which leads me to the crux of this email. Would you consider being a reference for me? And do you have any advice for this interview?"

I responded:

"Lol Excellent!  Good for you, Hermione.

Certainly I can be a reference (as a teacher) for you.
Probably the best advice I have for your interview is if you don't have the right answer, straight up tell them.  But then also tell them you'll have to answer (or know about whatever their asking you) tomorrow.  In other words, when you get home, you'll investigate it and get the answers.  
Bring a notepad to the interview and make notes about anything like that (so you look like you mean business).  Come with a couple of questions as well.  Research in advance anything in the job description you don't know about so you feel prepared.  Research the company a bit - know where their office is, ensure you can make it there on time, who are their current clients, some of the history, etc. 
Smile!  I don't know if you read my blog post about that, but smiling is HUGE.  If you can, try and get an interview somewhere else first to practice and get the jitters out (and maybe get a competing offer) 
Hope that helps!  Good luck!"

She replied:

"Thank you! I appreciate the reference and the advice.
I've been panicking a little, I really thought they would never call. I'm scrambling to get my portfolio site updated for the interview, as well as just get prepared in general. I do have a practicum lined up though, so no pressure...sort of."

In the end, Hermione got the internship.  She was nervous going into the internship because she didn't feel entirely qualified.  I told her not to worry and ask LOTS of questions.  She ended up successfully completed her internship and came out feeling better about it than she expected to.  It was a great lesson for her (and for me and all more students who I tell this story to) of how there are opportunities that you don't see in the job market.

Friday, July 6, 2018

Facebook Job Interview - Production Engineer

Facebook contacted me on LinkedIn recently looking to fill a 'Production Engineer' role.  I wasn't looking for job offers.  They reached out to me.   Apparently they had been doing this a bit though, targeting DevOps professionals as my buddy at the startup I was recently working with also got contacted.

Having Facebook reach out to you for a job opportunity?  Definitely intriguing.  One of the things on my 'IT Career Bucket List' would be to work at one of 'those' companies.  Google, Facebook, etc.  I thought if nothing else, it would be interesting to see where the hiring process went, since in the past they didn't actively recruit unknowns like me. 

I replied that I'd be interested to know more and so a phone interview was arranged with the Facebook technical recruiter.  At the appointed time (a couple days later) he called and I got the skinny on how things would potentially work.

Facebook logo - about my interview process at FaceBook

The phone interview was the first step.  After that I provide them with my resume which some of their technical leads would look at.  If they felt I was a fit based on my resume, there would be two remote technical screens - 1 specifically focussed on the Linux OS, and the other on a coding language of my choice.  They wanted someone who was very comfortable at both.  Over the phone they gave me some sample questions - basic linux commands - to give me an idea of what the screens would be like.  If I managed to gain their approval in the technical interviews, Facebook would then fly me to the location of my choice (Menlo Park or Seattle) for face to face talks - both technical and otherwise.  Following that, if I was still up to snuff, I'd get an offer.  Once hired, there would be a 6 week on site 'boot camp' where I'd get trained in all things Facebook and brought up to speed the technical ins and outs of the team I'd be working with.

Getting hired would require me to move to Seattle or Menlo Park.  Real estate at both locations is exorbitant - like ridiculous.  In the event that I was hired, Facebook would offer me a full relocation package, with the potential for a temporary housing situation for 3-4 months while we looked for a permanent residence.  I was told that many employees commute from communities with more reasonably housing prices using Facebook commuter busses that have complimentary drinks and WIFI.  Health and Dental would be covered 100% for myself and my dependents.  Wednesdays are an optional work-from-home day, and Facebook offers 21 days of vacation per year - although I neglected to confirm whether that was 21 business days (over 4 weeks), or 3 weeks all in.  I'm a Canadian, so I asked about a work Visa.  He replied that Facebook has an immigration/legal department that would arrange a TN1 Visa for me, in the event that I was hired.

As far as the job itself - Facebook's version of the Production Engineer role is essentially described here.  At the time I was talking with them, Facebook had 42 teams each responsible for a particular feature in their system (FB messenger, Ads, Newsfeeds, etc.).  These teams consist of 4-5 developers with an embedded production engineer.  Everyone on the team goes on call, one week at a time, so it ends up being a 4-5 week on-call rotation.

In the end, after viewing my resume they decided not to pursue the hiring process with me further at this time.  They had been interviewing 'a lot of strong candidates recently that they felt were a stronger match for their immediate needs.'  I can't say I was heartbroken.  It would have been a big move for us that would have put pressure on me personally and financially - not to mention having one dependent in university and one in high school.

I have been pondering what it was about my resume that flagged it to the hiring managers.  Was it the fact that I've moved jobs every 2-3 years, and they had concerns that I wouldn't last long at FB?  Perhaps it was because of my lack of focus in technology - moving contracts 2-3 years means learning lots of new technology and never getting a chance to really focus.  I've asked the FB recruiter - we'll see what he comes back with.  (You might be wondering why I've moved jobs every 2-3 years...  Its prudent for independent business contractors like myself to 'keep moving' from a Canadian taxation perspective.)

Saturday, September 9, 2017

My Path to the AWS Certified Solutions Architect - Associate Exam


Its been a while since I've tried to get any kind of industry certification. Life was busy. I was full-time consulting and had clients to support after work as well. Why would I even bother with all the study? Does a certification amount to much anymore? Lately, several things converged to motivate me to write this cert exam...

  1. My resume was getting stale. While I was able to keep a steady stream of contracts going, many of them were using older technologies that weren't current. This concerned me.
  2. I got a new gig at a start-up that gave me a hands-on opportunity to work with newer technology in the cloud. After setting up the infrastructure and CI/CD for several QA environments and geo-load balanced UAT and PRD environments in the Google Cloud Platform, a requirement for hosting our data in Canada became paramount. AWS had recently launched data centres in Canada, and so in May we decided to migrate our infrastructure there. Having successfully completed that migration I wondered how easy it would be for me to follow up with a Solutions Architect certification.
  3. A couple of guys at work got me turned on to udemy.com. Catching some sales I was able to get a couple of courses on AWS certification for $15 each (regularly they are over $150 each). One course was by ACloudGuru for the AWS SysOps certification. The other was by the Linux Academy for AWS Solutions Architect certification. These two courses gave me a good foundation for the material that is covered on the exam.
I originally did the SysOps course first, as I thought the that course was more in line with my job description at work. Finishing that course (leaving its practice test for later), I took a look at the practice questions on AWS here and felt like I'd come up short if I wrote the exam. That's when I decided to switch gears and do the Solution Architect course and exam. Most people online recommend doing that one first.

Incidentally, both courses helped me get a better grasp AWS Best Practices, and I was able to implement several improvements to our infrastructure at work because of what I learned. After another week and a half of going through the Solutions Architect course and reviewing the material, I felt more confident. I took the practice tests from both courses and passed them with good marks, so I thought I was ready. I scheduled my exam for the next week and also, for good measure, purchased a 20 question 'official practice exam'.

I had the week off during which I was scheduled to write my exam. Thursday was the big day. Tuesday morning I wrote the 'official practice exam' which was full of scenario based questions - and got 60% - A FAIL! Apparently the passing mark floats a bit between 62 and 66% (go figure). The practice exams in the courses seemed a bit easier. One of the things I didn't realize was how the different 'domains' for the exam were weighed:

In my 'official practice exam' I had scored:
1.0 Designing highly available, cost-efficient, fault-tolerant, scalable systems: 50%
2.0 Implementation/Deployment: 100%
3.0 Data Security: 75%
4.0 Troubleshooting: 50%

Clearly my 100% in implementation/deployment wasn't going to help me much with the domains balanced like that. I had some work to do!

A Cloud Guru has a forum specifically for the exam. I poured of that, specifically looking at posts where people who've written the exam share their experiences and what they wished they would have studied. I made notes of what I didn't know from those posts, and I also went over a lot of official AWS documentation for specific services (mostly FAQs and Best Practices) and took notes from them, too. And then I studied hard. 6 pages of 8 point font text. I kept adding things to those notes as well.

Thursday rolled around and I wrote the exam. I tried my best. The exam is 55 multiple choice questions and you have 80 minutes to complete it. Some of the questions I knew the answer, hands down. Others (like 'choose two correct answers') got a bit dicey. I chose the answers that made the most sense to me. I didn't feel like I was blind-sided by any of the questions, however there were definitely some things I (still) could have studied more (AD/User Federation types, for example). I finished answering the questions with about 15 minutes to spare. I had flagged some questions, so I went back and reviewed them and changed a couple of answers. When all was said and done, I passed and scored 72% - not great, but good enough for a certification. 

Here's how things panned out:
1.0 Designing highly available, cost-efficient, fault-tolerant, scalable systems: 75%
2.0 Implementation/Deployment: 80%
3.0 Data Security: 55%
4.0 Troubleshooting: 80%
I was quite happy with the improvements I'd made in the first and last domains.
Apparently there is quite a bit of content cross-over between the AWS Solution Architect - Associate exam and the other two associate certifications: SysOps and Developer. Potentially I could study a bit and write those fairly soon. However its $150 USD to write the exams, and I'm wondering what difference is in having the 3 certs verses just the 1. I'd be interested to hear your thoughts. Would it be work the extra $300 USD? Personally, I'm content to take a break from studying for now

Saturday, August 20, 2016

Google Cloud Platform & Stackdriver Monitoring - First Impressions

I've been working with the Google Cloud Platform at work for a little over a month now, and I'm getting comfortable with it.  I haven't done an exact price comparison with EC2/AWS, but apparently its a bit cheaper.  Based on my last 6 weeks of experience, I'm pretty happy with the pricing so far.  Static IP addresses are free, and their dashboard tells you where you can make hosting optimizations to save money which I think is great.  There's lots of good documentation, and the libraries for adding software are full featured and work well.  While GCP not be as full-featured as AWS, I find their menus and dashboards easier and more intuitive to navigate than AWS.

Hiccups I ran into:
- I can't dynamically add cpu or memory.  I've got to stop the instance and restart it.  Same with disk space.  I'd also like to be able to make some of my disks smaller, but I can't seem to do that either without some kind of reboot.
- If an instance has a secondary SSD drive,  it seems that there are issues with discovering it on a 'clone' or a restart after a memory or CPU change.  I have to log in through the back end (fortunately GCP provides an interface for this) and comment out the reference to the secondary drive in the /etc/fstab file to get it going again.  This seems buggy.

I recently spent a couple of days configuring Google Stackdriver Monitoring this week, and after a
couple of frustrations with not being sure how to get started, I installed an agent on a server and I was off to the races.  Installing agents is definitely the way to get going.  I found that it was easiest to create Uptime Monitors and associated Notifications at the instance level, on a server by server basis.  Doing it this way allowed me to group the Uptime Monitors to the instance, something I couldn't do when I created the Uptime Monitors outside of the instance.  It was simple to monitor external dependency servers that we don't own, but need running for our testing.  I integrated my Notifications with HipChat in a snap.  I also installed plugins for monitoring Postgres and MySql - these worked great so long as I had the user/role/permissions set correctly.  I'm super impressed with Stackdriver Monitoring, and will probably use it even if we have to switch over to host with AWS.

Our biggest roadblock with the Google Cloud Platform currently is they don't have a datacentre in Canada.  That could be a big selling feature for many of my clients, and because they don't have it, we may have to consider other options (like AWS/EC2, which is spinning up a new data centre in Montreal towards the end of 2016...)  Privacy matters!  Hope you're listening Google...

Thursday, April 7, 2016

EC2 and the AWS (Amazon Web Services) Free Tier - My First Experience

The Amazon Web Services Management Console
I had my first go-round with EC2 in AWS this last week in a real-life context.  I was teaching a class on Content Management Systems at SAIT, and I wanted my students to experience what its like to install WordPress and MySQL on a Linux VM.  I also wanted to personally get some experience with AWS, so I thought 'why not kill two birds with one stone?'

About 6 weeks ago I had purchased Amazon Web Services IN ACTION, written by Andreas and Michael Wittig and published by Manning.  It gave me a great primer on setting up my AWS account, my Billing Alert, and my first couple of VM's.  I leveraged that experience, crossed my fingers, spun up 18 VM's for my students, and hoped I wouldn't get charged a mint for having them running 24/7 for a few days.  It was a Friday around 1pm when I created them and gave them to my students to use.  Imagine my surprise when I checked my billing page in AWS on Monday and discovered they had only charged me $2.81!

I had clearly reached some kind of threshold as after that day I got charged, on average, about $8/day - for all 18 VM's.  They were charging me 2 cents per hour per VM, and some small data transfers.  Granted, I used a 'T1 Micro' - with 1 CPU, .613 Gib memory and 8 GiB storage.  Still, I was quite happy.

On the last day of class, I split the students up into two teams and gave them large web projects to do, and spun up a 'T1 Micro' for each team.  I gave them some scripts they could run to create some swap memory if they needed it.  They ran those scripts right away, and within an hour (with 9 people uploading files and content into those systems continuously) those T1 Micros CPUs pinned.  So I quickly imaged them over lunch and spun up an 'M3 Large' VMs (2 CPU's, 7.5 GiB memory) for each team and threw their images on.  I ran into one issue spinning up the new team VM's - I had to spin down the original Team VM's first before I could start the new ones because there is a limit/quota of 20 VMs on the 'free-tier' in AWS.   Aside from that and the changed IP addresses, the transition was seamless.  I was a happy camper and now that the students had responsive VM's, so were they.

My total bill for the week - $38!  A colleague pointed out that I probably could have added auto-scaling to those team VM's and increase the CPU and memory in place, without losing the current IP's.  He's right, I probably could have, but I didn't have the experience and didn't want to waste class time (and potentially lose the student's work) by trying something I didn't know how to do.  All in all, I was very impressed with my first run of EC2 in AWS.  It was very reasonable, responsive, and easy to use.  I'd definitely do it again.


Monday, March 28, 2016

The Perils and Pitfalls of OpenSource Software

Open-Source Software is the backbone of most of the internet.  Seriously.  For more than a decade, small business websites and main-stream web applications have used open-source software to develop their solutions.  Web projects (and frankly, most software projects in general) that don't have a dependency of some kind on an open-source component are the RARE exception to the rule.

Should we be concerned about this?  I think so.  Here's why:
  1. Coders aren't implementing open-source code properly.  In my experience, any open-source code dependencies should be referenced locally.  Many coders fail to do this and reference external code libraries in their code.  What happens when that external 'point of origin' has a DNS issue, or is hit with a DDOS attack, or is just taken down?  Your site can break.

    Case in point - check out this article on msn.ca 'How One Programmer Broker the Internet'  In a nutshell, one open-source programmer got frustrated with a company over trademark name issue.  This developer's open-source project had the same name as a messaging app from a company in Canada.  He ended up retaliating by removing his project from the web.  It turned out his project had been leveraged by millions of coders the world over for their websites, and once his code was removed, their websites displayed this error:
             npm ERR! 404 'left-pad' is not in the npm registry

    I believe if developers had downloaded the NPM javascript libraries and referenced them locally on their servers, they wouldn't have run into this issue (as they'd have a local copy of the open-source code).

    Another case in point - I worked on a project a number of years ago that had dependencies on jakarta.org - an open source site at the time that was hosting a bunch of libraries and schemas for large open-source projects like Struts and Spring.  Some of the code in those projects had linked references to schemas (rules) that pointed to http://jakarta.org.  Unfortunately we didn't think of changing those links and one day jakarta.org went down....  and with it went our web site, because it couldn't reference those simple schemas off of jakarta.org.  After everything recovered we quickly downloaded those schemas (and any other external references we found) and referenced them locally.
  2. Security  Do you know what is in that open-source system you're using?  Over 60 million web sites have been developed on an open-source content management system called WordPress.  Because it is open source, everyone can see the code - all the code - that it is built with.  This could potentially allow hackers to see weaknesses in the system.  However, WordPress has pretty strict code review and auditing in place to ensure that this kind of thing doesn't happen.  They also patch any issues that are found quickly, and release those patches to everyone.  The question then becomes:  Does your website administrator patch your CMS?

    Another issue I ran into related to security was with a different Content Management System.  I discovered an undocumented 'back-door' buried in the configuration that gave anyone who knew about it administrative access to the system (allowing them to log into the CMS as an Administrator and giving them the power to delete the entire site if they new what they were doing.  Some time later, I found out that some developers who had used this CMS weren't aware of that back door and left it open.  I informed them about it, and they quickly (and nervously) slammed it shut.   Get familiar with the code you are implementing!
  3. Code Bloat (importing a bunch of libraries for one simple bit of functionality)  Sometimes developers will download a large open-source library to take advantage of a sub-set of functionality to help save time.  Unfortunately, this can lead to code bloat - your application is running slow because its loading up a monster library to take advantage of a small piece of functionality.
  4. Support (or the lack of it)  Developers need to be discerning when they decide to use an open-source library.  There are vast numbers of open-source projects out there, but one needs to be wise about which one to use.  Some simple guidelines to choosing an open-source project are:
    • How many downloads (implementations) does it have?  The more, the better, because that means its popular and likely reviewed more.
    • Is there good support for it?  In other words, if you run into issues or errors trying to use it, will it be easy to find a solution for your issue in forums or from the creators?  
    • Is it well documented?  If the documentation is thorough, or if there have been published books written about the project, you're likely in good hands.
    • Is it easy to implement?  You don't want to waste your time trying to get your new open-source implementation up and working.  The facilities and resources are out there for project owners to provide either the documentation or VM snapshots (or what have you) to make setting up your own implementation quick and easy.
    • How long has it been around?  Developers should wait to implement open-source projects with a short history.  Bleeding-edge isn't always the cutting edge.  Wait for projects to gain a critical mass in the industry before implementing them if you can help it.


Friday, March 18, 2016

Random Thoughts About the DevOps Movement in IT - Part 2

The Phoenix Project (image of the book on the left) was recommended to me over a year ago as a good read about DevOps.  It emphasizes the benefits of Dev Ops to management, shareholders, and a company.   Impressions I had of the book were:

  • Its a great, fun read, however it skims over the difficulties and trials of the actual automation of the technical systems and software development process - where the rubber really hits the road.  It seems to practically romanticize the idea of automation in IT a bit - like DevOps is the goose that will lay your golden egg.  Unfortunately, there's significantly more work to get that golden egg in my experience.
  • It also glosses over how to get the Security team on board with what DevOps wants to do.  In many companies, the Security team holds the trump card and if they decide to change all your certs from 128 bit encryption to 2048 bit encryption (don't laugh, I've seen it happen and we had to regenerate certs for all applicable servers in all envs).  Their wish is your command unless you can convince someone influential that that kind of encryption is overkill in a non-prod environment.

Finding and Cracking That Golden Egg
Hurdles I've encountered enroute to the DevOps 'golden egg' are:
  • Silo-ed Application Projects.  
    • Lack of consistent naming conventions for deployment artifacts and build tags/versions is majorly detrimental to the automation process. 
    • Lack of understanding of the dependencies between application projects (API contracts, or library dependencies).  If you don't understand your dependencies between applications, they may not compile corectly, or they may not communicate properly with each other.
    • Lack of consistent development methodology and culture between teams.  If you have one team that doesn't get behind the new culture, and they insist on manually deploying and 'tweaking' their code/artifacts after deployment, they risk the entire release.  Getting everyone on board with culture is challenging.  
  • Overusing a tool.  When you have a hammer (a good DevOps tool), everything is a nail.  Generally, many of the DevOps tools out there all have their niche.  You could potentially use Chef, Puppet, or Ansible to do all your provisioning and automation.  Is that the best solution though?  I'm inclined to say no.  How much hacking and tweaking are you having to do to get all your automation to work with that one tool?  Use the tools for what they are best at, what they were originally made for.  Many of these tools have Open Source licenses and new functionality is being added to them all the time.  While it might seem that the new functionality is turning your favourite tool into the 'one tool to rule them all',  you might have a big headache getting there, trying to push a square peg in a round hole.  
  • Lack of version control.  Everything must be stored in a repository.  The stuff that isn't will bite you.  VM Templates, DB baselines, your automation config - it all needs to go in there.
  • DB Management.  All automated DB scripts should be re-runnable and stored in a repo.  
  • Security.  How to satisfy the Security Team?  How to automate certificate generation when they are requiring you to use signed certs?  How to manage all those passwords in your configuration?
  • Edge Cases.  Any significantly sized enterprise is going to have/require environments that land outside your standard cookie cutter automation.  Causes I've seen for this are:
    • Billing cycles - we needed a block of environments that were flexible with moving time so QA didn't have to wait 30 days to test the next set of bills and invoices.  
    • Stubbing vs. Real Integration - Depending on the complexity of integration testing required, there may be many variations of how your integration is set up in various environments - what end-points are mocked vs. which ones point to some 'real' service. 
    • New Changes/Requirements - Perhaps new functionality requires new servers or services.  This can make your support environments look different that your development environments.
  • Licensing issues.  When everything is automated, it can be easy to loose track of how many environments you have 'active' versus how many you are licensed for.  License compliance can be a huge issue with automation - check this interesting post out 'Running Java on Docker?  You're Breaking the law!'
  • Downstream Dependencies   This is where the Ops of DevOps comes into play.  Any downstream dependencies that your automated system might have need to be monitored and understood.  You can't meet your SLA with your client if your downstream dependencies can't meet that same SLA.  Important systems to consider here are: LDAP, DNS, Network, your ISP, and other integration points.
  • YAML files.  Yes, perhaps they are more terse than storing your deployment config in XML.  However, I'm at a bit of a loss to see how they are better than a well named and formatted CSV file.  Sure, you can 'see' and manage a hierarchy with them.  But you can do the same in a CSV file with the proper taxonomy.  YAML files utilize more processing power to parse and have extraneous lines in them because their trying to manage (and give a visual representation of) the hierarchy of the properties.  I've seen YAML files where these extra lines with no values account for a significant percentage of the total lines in the file, making the file more difficult to maintain and prone to fat fingers.  Several major DevOps tools use these files and I really can't see a good reason why except that they were the 'new, cool thing to do.'

Wednesday, March 16, 2016

Random Thoughts About the DevOps Movement in IT - Part 1

DevOps related books I have in my library currently
Almost 5 years before DevOps was given the name it has today, I helped implement a modern DevOps solution at a company I was consulting with.  While our enterprise wasn't huge, and we didn't have a lot of integration (compared to some) it was a pragmatic solution that used Kickstart and Ant with a PHP/Mysql database.  We had Continuous Integration (CI) and Continuous Delivery (CD) - members of our QA team could select an environment, select a build, press a button, and the deployment happened.  Our environments at the time consisted of JBoss containers on Unix machines, and IIS web-apps and services on Windows machines.

Software development methodologies and practices have evolved a lot since then.  It's been a challenge to keep up.  I'm a pragmatic guy, and I thought one could get pretty much all the functionality needed for an automated build and deployment stack just using Ant, CruiseControl, Kickstart, VMWare, and some helpful API's.  Clearly I was wrong.

Cause for Pause
Without a doubt, the DevOps movement has captured the imagination of many a software developer.  Just look at all the tools out there now!  Chef, Puppet, Bamboo, RunDeck, Octopus, Docker, TeamCity, Jenkins, RubiconRed, Ansible, Vagrant, Gradle, Grails, AWS...  I could go on.  I'm beginning to question whether or not the polarization and proliferation of these tools has been helpful.  I find many companies looking to fill a Devops role are asking for resources who have experience with the specific tool stack they are using.  That must make human resource managers pull their hair out.  Even personally, I'm concerned about hitching my wagon to the wrong horse.  If I decide to accept a position with a company who is using a Chef/Docker stack (for example), but the industry decides that Ansible/Bamboo is the holy grail,  have I committed professional suicide?  Probably not, but it does make one consider job opportunities carefully.

Looking Forward
This P and P (proliferation and polarization) of DevOps tools makes me wonder what the future is going to look like for the DevOps movement.  Consider what Microsoft did with C#.  Instead of continuing to diverge and go their own way with the replacement for VB, they created a language with a syntax that essentially brought the software development industry back together (in a way).  Brilliant move.  Java developers quickly ported their favourite tools/frameworks over to C# and suddenly developing with Microsoft tools was cool again.  It would sure be nice if something like that happened with the DevOps tools.

Something else to consider for the future...  so far in my experience with overseas, outsourced teams, they have not yet embraced the DevOps movement.  As a result, they aren't as efficient or as competitive as they could be.  When they do jump on the DevOps bandwagon and truly tap into it's potential, it could be a game changer for people like me....

Tools to Watch
  • Perhaps Amazon is the 'new Microsoft' with its expansive and ever expanding AWS tool stack?  There's definitely some momentum and smart thinking going on there.  They appear to have automated provisioning all wrapped up in a bow. 
  • Atlassian is another company (I didn't realize they were based out of Australia) that is putting together DevOps tools that have a lot of momentum in the industry.  Their niche is collaboration tools.
  • Puppet and Chef both have a strong foothold in the DevOps community.  Both of them are adding new features all the time, enabling them to automate the deployment and provisioning of more products and systems.  Many people use Puppet and Chef in conjunction with RunDeck or Docker to get all the automation they are looking for.





Friday, March 4, 2016

Teaching an IT class (JavaScript, XML, C#, VB, CMS/PHP)

I've taught part-time at SAIT for 11 years now.  I've had more than 30 different classes with 5 different courses.  Recently I've been challenged by a colleague (also a part-time teacher here at SAIT) to make my courses better and more interactive.  While I thought my courses were pretty interactive to begin with, his ideas have definitely pushed me further in a good way.

All the courses I've taught have been introduction to different code languages or working with code:
Visual Basic, C#, XML, Javascript, PHP, and Content Management Systems like WordPress and Drupal.  Up until recently, my courses had been - at a high level - structured as follows:  I demonstrate a new concept/code idea to the class and I have them follow along.  Once they are reasonably comfortable with implementing it (perhaps an exercise or two of following along like this) I give them an exercise that they have to try and do themselves.  Sometimes I'll add a twist to make them think a bit.  For the most part this has worked great.  Its looked like this (an example with XML):
  • Intro to the Language and syntax - I'll go through several power point slides that discuss: 
    • What does well formed XML mean?  
    • How do you ensure your XML is well formed?  
    • What are the keys/rules to well formed XML? 
    • I'll ask the class questions as we go along and get them to answer.
  • We'll code a simple XML file together.  I'll make intentional mistakes along the way and see if the class is paying attention, or just call them out and ask 'Do you think this is right?'  I will also put comments in my code so the students can refer to them later.  Once done, I'll give this complete exercise to the students so they have it to refer to.
  • Then I'll give them an XML file that is not well-formed (completely broken) and get them to try and fix it.  
    • After about 15 minutes, I'll try and fix it myself in front of them, getting them to help me.
    • Once that's done with my comments, I'll again put it in a shared folder for the students to grab so they can refer to it later.
  • I'll give them another broken XML file and get them to fix it. 
    • Again, after about 15 minutes, I'll try and fix it myself in front of them, getting them to help me.
    • Once that's done with my comments, I'll again put it in a shared folder for the students to grab so they can refer to it later.
  • Then, next concept.... and repeat.
This has worked quite well.  I've followed this process with all my courses and it seems to keep students engaged and helps them grasp new concepts.

As I mentioned earlier, I've been challenged by a colleague to make my courses better, even more collaborative, and (hopefully) gets the students more return on their tuition investment.  Here's what I'm changing and improving:
  • Using a Repo.  Instead of using a shared folder to give files to students and have them turn in assignments, I'm using BitBucket with SourceTree (and forcing them to use it too).  This is great practice for them for the industry, and so far has worked really well.  My SourceTree client does get bogged down at times (with over 30 repositories) but it hasn't been a major deal.
  • Commenting. I continue putting lots of comments in my code for the class to refer to.  They really like this and find it helpful when they review after class.  I'll definitely keep doing this
  • Error Challenge.  I tend to fat-finger a bit when I'm coding in front of the class (even when I'm working off of a printed copy of a complete exercise).  Sometimes these errors are hard to catch and I waste some class time while I try and figure out what I did.  So instead of waste our collective time and to help them pay attention, I instituted a little game.  If a student catches an error I make in my code, they get a point.  Points accrue for the duration of the course.  The student at the end of the course with the most points gets a free Starbuck's gift card (or something like it).  This worked really well, and I'm definitely going to do it again.
  • Student Journal. Student feedback is key for me.  I teach in 'Fast-Track' programs where the students know that they are going to get inundated hard and fast with new concepts all the time.  Ensuring I'm not bull-dozing them with new stuff is key, so having a feedback loop helps.  Up to this point I've just been asking them in class 'How's the speed - am I going to fast?'  But my colleague suggested getting the students to journal and having them  commit that to their repo so I can see it as another feedback loop.  I did this in my last course and it worked quite well.  I don't have them journal every day - just once a week or so, but the feedback I got from them individually was great. 
  • Group Projects.  Out in the industry, most of these students are going to be working on teams, so having them do exercises in class as teams seems logical.  It turns out that students, for the most part, really like having group projects to work on.  They can collaborate together and learn from each other.  It also pushes them to compete a little which makes for an even better result.  I have them create a team repo for the project(s), and this also gives them experience with some BitBucket administration (branching and merging, etc.) which also prepares them for the industry.

Sunday, January 24, 2016

IOException - Disk Quota Exceeded error - yet there's lots of disk space


I use a Mac for work.  I've run into an interesting issue with making backups of client web sites and zip files with my Mac...  It seems that whenever I've zipped a file and backed it up to my Mac, I get these wonderful __MACOSX folders - that's 2 underscores at the beginning of the folder name injected into the folder mix.

Restoring a site from one of those backups on a linux machine can have adverse consequences because of those 'injected' __MACOSX folders.  It seems that certain flavours of linux really don't like that folder name and as a result will not write a single file to the machine any more.... at all.  Instead, it throws an a 'Disk Quota Exceeded' error.  You can try touching a file, you can try unzipping some other files - whatever you do, you'll get that error.   The error didn't go away for me until I deleted all __MACOSX folders that got copied up to my machine in the zip file.  Once I did that - it was all good.

The interesting thing about this error was I could have tons of disk space still on the machine, and yet it would still throw that error until I removed those offending __MACOSX folders.

There are various commands out there for removing these folders from zip files created on a Mac... something like this:
zip -d your-archive.zip "__MACOSX*"

Monday, January 18, 2016

Get Hired as an IT Newbie Part 2 - Soft Skills

I participated in a focus group this past week at SAIT where we discussed their 'Web Developer Fast-Track Program' and its relevance to the industry.  The goal behind the meeting was how can we make the course more relevant to the changes and continual evolution of technology that is happening in the industry so students are more prepared for positions when they leave.  The discussion was engaging and took some interesting turns.  Here's what I took away related to Soft Skills:

Soft Skills
Employers like student hires to be reasonably technically adept.  Their technical expectations aren't in the stratosphere.  However, they want their student hires to have more than just technical ability.  Students should be well versed in soft skills.

Soft skills is (to some degree I'm sure) an ambiguous subject for many students.  I can hear you thinking 'What do people mean when they say that??'  In the context of our discussion  yesterday,  it meant prospective IT employers in the focus group yesterday were looking for students to have, in addition to technical skills:

  • An ability to communicate well with a client - verbally, directly in front of them or over the phone.  Also having the discernment to know when to send an email versus when to talk verbally.  Poorly written emails are notorious for communicating the wrong things - particularly negative feelings where there were none.
  • An ability to know what good work is.  Does the design and function of a web site fit the purpose/company it was designed for?  For example, does a sales/marketing website have lots of good calls to action?  Does the site feel right?  Would 'Joe's Mom' know how to navigate and use the site.
  • A can-do attitude.  Students will likely get the soul-crushing repetitive work that senior developers/designers don't want to do (or don't have time to do).  Students should accept this and be prepared to do it with gusto.  If you're in an interview and you don't have a clue about a technology they asked you about, reply: "I don't know about that, but I will know about it tomorrow" and mean it.
  • An ability and desire to collaborate and work in a team setting.  Its a fact of the industry - if you want to work in an agency, you're going to need to be able to work on a team and collaborate with team members. 
Also discussed was: What habits should students have (or start to form if they don't have them)?
  • Students coming out of school should love learning and know how to learn things/solve problems on their own.
  • Students should be able to receive constructive criticism without responding with excuses about bad instructors, poor course material, or the speed at which things were taught.  You don't get much chance for excuses on the job.  Employers are looking to 'break even' economically on their student resource investment within 4 months.  

Saturday, January 16, 2016

Get Hired as an IT Newbie Part 1 - Have a Good Portfolio!

I participated in a focus group this past week at SAIT where we discussed their 'Web Developer Fast Track Program' and its relevance to the industry.  The goal behind the meeting was how can we make the course more relevant to the changes and continual evolution of technology that is happening in the industry so students are more prepared for positions when they leave.  The discussion was engaging and took some interesting turns.  Here's what I took away related to portfolios:

Student Portfolios
15 years ago I had gone through a similar IT fast-track program and we had to do portfolios of our work for potential employers.  I wasn't aware of a single employer who looked at my portfolio and as
a result I haven't placed much emphasis on it in my classes.  BIG MISTAKE.

It turns out that prospective employers for SAIT students DO look at their online portfolios.... and generally weren't super impressed.  Here's why:
  • Industry attendees said that a student's portfolio should reflect the position or career track that the students are looking to get into.  There should be evidence in the portfolio that the student tried on their own to investigate, explore, and work with technologies and code that interests them.  Posting student projects (with every student having very similar projects) doesn't help anyone in making a hiring decision.
  • Prospective employers are not just interested in the technologies that students used, but also: 
    • Thought processes that the students went through in completing their projects - why they chose to make certain decisions (for example, use a canned WordPress theme instead of design/develop their own)
    • Challenges students encountered in design and development and how they overcame those problems in their journey to the completed project - Employers are interested in soft skills like perseverance,  the ability to google a problem and uncover a solution on your own, the ability to asking for help when you are stumped (not before you've tried googling the problem)... etc.
    • Is the student willing to go the extra mile, put in some extra effort and explore interesting tools and technologies outside the classroom?  The industry was quite clear that this was one of the main things they were looking for  - a motivated individual with the right attitude.  This kind of motivation should be clearly seen in a student's portfolio.
There was also some suggestions during the focus group that students should be given the opportunity to provide constructive feedback to their peers on their portfolios.  A good critique is a gift.  Can students accept constructive feedback and use it to improve?  Could providing peer critiques give them a better perspective and more experience with what is good (design, code, functionality, etc.) versus what is not good.  We thought so.

Finally, here's a couple of links to student portfolios that 'make the grade' so to speak, in my humble opinion.  One was a student of mine this past semester, the other is a current student at the University of Calgary:
- Gary S. Jennings (former SAIT student)
- Carrie Mah (U of C student)

Sunday, March 1, 2015

Books read to date in 2015

One of the advantages of a good commute is having the chance to get some good reading done.  I've taken a different tact the last few months in my technical reading in that I'm choosing books that I hope I'll enjoy.  For the most part, its been a good experience.  Here's what I've managed to get through so far this year:

The Cuckoo's Egg by Cliff Stoll.  This has been the only 'reread' for me so far this year.  If you're in IT, you might this this book and enjoyable read.  Cliff recounts his adventures back in the late 1980's chasing a hacker internationally on the ethernet through university, military, and other government networks and computers.  If you consider yourself a geek and you haven't read this book, you really should.



The Phoenix Project by Gene Kim, Kevin Behr, and George Spafford.  The team leads at my current contract recommended I read this book.  It's marketed as 'A novel about IT, DevOps, and helping your business win.'  I enjoyed this book (and probably most people do) because it was easy to associate characters in the book with real people I've worked with.  On top of that, the problems the characters confront in their business are taken from 'last month' in many IT shops around the world.  Readers can relate to this book.  In some ways, I could say I've lived this book about 8 years ago, and I'm trying to help implement those solutions again where ever I end up working.


Social Engineering in IT Security by Sharon Conheady.  I enjoyed this book more than I thought I would have.  Sharon doesn't only give a great overview of the tools, tactics and techniques of IT related social engineering, she also gives it a historical context and augments the 'theory' with real life experiences.




Cyber Warfare by Jeffrey Carr.  I found this book to be something of a disappointment.  Even though it was a second edition (2011) I didn't find many updates.  I believe its intended target audience is CIOs and it seemed to me that the author included more legal documents and previously published papers in this book it needed.  I seriously considered returning this book and getting my money back.



Callings by Gregg Levoy.  I really enjoyed reading this book, perhaps because I'm very interested in the topic.  Certain readers might find it a bit academic, but there are many interesting, true stories that keep the narrative engaging.  This is NOT an IT book - rather it more about 'finding and following and authentic life.'  I originally got this book from the library on a whim, but after reading it I bought a copy.  Beware - some editions of this book (like the one I got from the library) have the last 20 pages out of order.  Check the page order from 300 on before you buy your copy.

Wednesday, January 14, 2015

Enterprise Architecture in Calgary

I went to a CIPS (Canadian Information Processing Society) lunch on Tuesday and the topic of discussion was 'The State of Enterprise Architecture'.  The talk was given by Martin Mysyk, the curator of www.enterprisearchitecture.ca.  Actually, it was more of a discussion lead by Martin.  This was appreciated as there were many practicing architects in the room and it was good to hear the variety of perspectives, challenges, and concerns that these fellows face.

Martin began by posing the question 'What do Enterprise Architects Do?'  Someone responded from the back of the room, 'We slow projects down!'  Everyone laughed and then we got more serious.  Other answers were offered:  'we provide strategic vision, influence, and guidance to the enterprise',  someone read a Wikipedia definition for EA's.  In the end, Martin compared EA's with municipal planners.  I thought this was interesting and decided to google it, and I found this page:  Enterprise Architecture Analogies  Some interesting food for thought.

A slide was displayed about EA As a Career and a short discussion ensued on how does one get their foot in the door (certification, promotion from BA role, promotion from SA role, etc).  EA's require a delicate balance of technical and business expertise, combined with superb soft skills like communication and effective discussion leadership.  Martin recommended a book called the Zoom Factor by Sharon Evans.  I might have to pick this one up as it looks like it might be interesting.

Warming to his subject, Martin moved into the EA Practice.  He reviewed several different flavours of EA methodologies out there today, commenting that TOGAF seemed to be the current favourite.  A slide was displayed with EA Practice at the top, followed by four boxes (focuses of EA?):  Business, Data, Applications, and Technology.  I was surprised that Security didn't count among them, and it didn't seem to arise in the discussion as much as I thought it would.  If TOGAF suggests that a Business architecture is important, doesn't a Security architecture merit some consideration?  Perhaps I missed it in the diagram...

Some important considerations for EA's that came up in the discussion:

  • We need to consider growth versus optimization in Calgary right now (considering the price of oil)
  • Who in the organization should benefit from EA?   Is I.T. really the right place for every organization's EA practice? Some companies have their EA practice reporting to the Legal dept.  others report to Operations, or even the CEO.  
  • Perhaps the question is 'Where can the EA practice provide the most value in organization?'  Answers to these questions will be different for each situation.  Depending on the answers, one may have to reconsider who the EA practice reports to...
And finally our discussion wrapped up with the challenges that EA's face:
  • We need to prove our value with a small budget.  How do we do that?
  • How do we influence the decision makers without seeming heavy handed?
  • How do we provide guidance and structure to development and growth without seeming heavy handed?
I'd love to have some comments!


Tuesday, July 23, 2013

A couple more business ideas...

I think one could (somehow) put together a company that offers to host data for clients for free, in return for being able to datamine their data.  Obviously there would be privacy rules to consider and content with.  However, my intuition is telling me there are still opportunities in this area that are yet untapped.

My other idea would be to build light fixtures that incorporate a similar 'fan' model to the Dyson fan - in other words, no blades as part of the light fixture.  Rather, there would be a main hub that houses and hides the working fan.  The fan would blow the air at high pressure through slots in the light fixture.  The slots could be anywhere on the light fixture.  The fixture itself could potentially even move....

Based on the cost of these 'blade' fans, I'm sure a bundle could be made on similar light fixtures.

Adding another idea to this (Aug. 4).  I'm very surprised I don't see more rechargeable (hair) blow driers out there.  Along with that, hair clippers.  When I'm getting my hair cut, the stylist is always messing around because the cord isn't long enough, or the cord has knocked something over or gotten twisted or caught on something...   Isn't that what building contractors wanted cordless drills for?  They've been around for years. 

Wednesday, April 3, 2013

The Value of IT Certifications and Skills

After I've told my students at school how I finally managed to get a full time job in IT, they've asked me what I thought was the most important thing I did to get that job.  I've always answered 'getting my SCJP certification' (Sun Certified Java Programmer). It's been over 10 years since I started working at the job.  While I still maintain certifications are important for developers who are trying to break into the industry, I've become less enthusiastic about them and what they mean now that I have more experience. 

Having said that, I ran into a couple of articles today that I thought were very interesting:
15 IT Certifications That Can Get You a $100,000 Salary
30 Tech Skills That Can Get You a $100,000 Salary

Granted, the way these articles are written is 'just a little bit of' hype.  These skills and/or certifications aren't going to net you a 100k salary all on their own.  They need to be married with a resume that has several years of IT experience and related, well developed soft skills.

Monday, May 21, 2012

IT New Hire/Contract? Thoughts on IT On-Boarding

Being a contractor, I get to see and experience the on-boarding practice of different companies in different industries on a regular basis. This gives me an opportunity to provide a 'value-add' for the companies that I'm contracted with. Because of my fresh perspective, I can more easily pinpoint areas of opportunity for improvement in the first week on the job.

One thing that has always surprised me, and remained fairly consistent in all of the companies that I've worked with is that they don't seem to have the infrastructure ready for new hires.  I always seem to spend the first week doing one of the following:
  • waiting for a computer or even a desk to sit at.
  • waiting for access to my computer or to files, folders, applications, servers, etc.
  • setting up/configuring my computer - installing required applications, setting up internal website bookmarks in browsers, etc.
  • waiting for licenses for the applications I need to run to do my job.
I find this lack of organization dumbfounding.  HR knows that a new hire is getting on-boarded weeks in advance.  Similarly, the team/department should know as well.  I don't think there are any real good excuses for having new hires sit on their hands for the first week while everyone gets organized.  Bringing new hires on board should be a Standard Operating Procedure.

While I'm waiting for access or a license, I busy myself with looking for things that people with history at the company would view as the 'status-quo', but are really issues that could/should be considered problematic or improvable.

Opportunities like:
  • Single Points of Failure.  (email server(s), monitoring server(s), shared drive(s), VPN access, etc)
  • Lack of network segregation between environments (can I ping/telnet from a dev or staging server to production?)
  • Misnamed or legacy names for servers or monitoring notifications that are understood by employees but a source of confusions for newbies
  • Lack of templates for common documentation
  • Lack of exportable outlook email rules that can be shared with the team
  • Lack of exportable browser favorites that can be shared with the team
  • Consolidating and organizing documentation (critical documentation in 12 different folders scattered through different directories is a recipe for confusion)
  • Lack of Automation - continuous integration through to automating deployment.  Automating monitoring
  • Preventing Fat-Finger errors (adding confirmation messages in start scripts)
Granted, not all of these will be applicable for every IT position.  Also, share your findings with your new supervisor and get his blessing before going and changing everything.  Suggest with encouragement, don't criticize! You don't to come across like your trying to relieve someone of their position.  You are there to make things better for everyone.

So, next time you are on-boarded for a new IT position, keep these things in mind.  Hopefully they can help both you and your new company.  Also, if anyone has some other suggestions of things to look for, please add a comment!

Tuesday, March 13, 2012

business idea - Muffin wrappers

Make edible muffin wrappers - My wife had a problem with some new muffin wrappers she tried sticking to the muffin.  Why can't they make edible ones?  If you get the texture just right, it could change the muffin/cupcake industry....

In fact, making products out of 'by-products' is good business.  Henry Ford did it with waste wood from his Model T's.  He turned it into charcoal bricks that ended up being used for fire fuel.  This side product eventually turned into a company called Kingsford Charcoal.  If you've ever had a charcoal BBQ, I'm sure you've bought a bag of their briquets.

Monday, July 18, 2011

Dell Latitude Diagnostic Error Code 2000-0142 Hard Drive Failure

'Apparently' the hard drive on my laptop failed yesterday.  Out of the blue I got a couple of popup windows complaining about some dll not being able to load, and then the windows blue screen of death.  I thought maybe my machine got too hot, but I tried rebooting after letting it cool off and it didn't do anything after the Dell splash page during the boot.  So I ran a BIOS diagnostic and got the following error:
Msg: Error Code 2000-0142
Msg: Unit 1: Drive Self Test
Failed Status Byte=77
I called Dell and they basically confirmed what the diagnostic test said - my hard drive was fried.

Not wanting to give up on my hard drive that easily, I googled to see what I could do.  I discovered I could try to repair windows using the Windows Recovery Console.  So I found my Windows CD and booted from it and went to the Repair option.  Once I had entered the Admin password, I got a command prompt and after a bit of searching around I discovered that most of my hard drive seemed to be intact.  This, however, was still in the dos prompt of the Windows Recovery Console.  I couldn't back up my files to a USB from there.

I tried chkdsk, chkdsk /p, and chkdsk /r - those all worked successfully.  Yet my machine still wouldn't book.  So I preused the list of commands available to the Windows Recovery Console (here) and tried a few more - bootcfg, fixboot, and fixmbr.  In the end, it was fixmbr (which repairs the master boot record on the specified disk) that ended up working for me.

My laptop is 4 years old now.  I removed the battery a couple of months ago, and with problems like this starting to happen, I going to start shopping for a new one.  Hopefully this info helps someone out there though.