2012/10/04

Telco Customer Service Madness: Case Study

Will Telstra, as it is now, survive to see the NBN contracts end in 35 years?
My view: It won't, not in its current form because of multiple failures within the Organisation.

Below is a case study that Telstra should deeply investigate as it encapsulates most of their challenges/deficiencies and could be used as an on-going Reference for Change, but why would they?

In an ideal world, the centre of the study would have these outcomes:
  • A personal meeting with the Head of Telstra for the State.
  • An apology from him, a guarantee it would never happen again and his personal phone number if further problems arose.
  • A desk audit of all records for their services and accounts to correct all errors.
  • A written account of:
    • Exactly what went wrong,
    • Why it couldn't be fixed, and
    • Why it won't recur.
  • An offer of compensation for the non-supply of service, for the hours of customer time wasted on the phone and waiting and an ex-gratia payment for the "pain and suffering" caused.
Case Study

The facts of the case study are:
  • Customer, 'A', has on their account multiple individuals, multiple service addresses, and multiple services for each individual and service address (mobiles, landlines, ADSL, Cable TV, Cable Internet, ...).
    • Whilst these are all domestic services, Telstra regularly deals with this complexity for SME's.
    • They are a "high-value" Telstra customer. This seemed irrelevant in the process.
    • Unsure if all individuals and services are billed together or by separate, linked accounts.
  • 'A' is also a Telstra shareholder, which seems to have been irrelevant in the process.
  • 'A' is highly educated, has run businesses and is well conversant with modern PC's and networking, relying on it for work and private life.
    • There are multiple family members who are quite I.T. literate and provide in-home I.T. support and troubleshooting.
  • A new Cable Internet service was ordered by 'C' in June. (date?)
    • The modem was never delivered.
    • When queried at the Telstra shop, customers were advised "the order had been cancelled".
    • The customers had not cancelled the order, nor been advised of that action.
  • 'A' had a working Cable Internet service that then became intermittent. It met their needs and wasn't reported as a fault due to very poor past customer experiences.
    • "Not wholly broken, don't tempt fate" was the reasoning.
  • 'B', another of the service holders, took it on themselves to report the fault to Telstra.
  • The first technician attended on 23rd-August, intending to change the cable modem.
    • They were unable to rectify the fault, did not replace the cable modem as it was serviceable and left saying "there is an error", which at some point changed to "an activation error".
    • The replacement cable modem was left on-site, unconnected.
    • 'A' was told the install failed because of "Error Code CCP0012",  and Tech suggested that the system “thought” there was already a modem on order.
    • Technician advised 'A' to call the general BigPond Enquires number (137 663), quote the Error Code, and the fault would be fixed.
  • Multiple technician attendances were booked:
    • Technician did not attend, did not phone customer. More than once? (date?)
    • Technician sent to wrong address, an old service address on the account. (17-Sep-2012).
    • Technician 'M' attended (19-Sep-2012), gave customer personal contact number and spent considerable time on-site and continued to work at resolving the fault.
      • Possibly instrumental and worthy of commendation.
    • 'M' followed-up a week later (25-Sep-2012) saying:
      • TRG (Technical Response Group?) were aware of the problem,
      • other customers (in the area, state, nationally?) were affected and
      • TRG didn't know when or if the Error Code could/would be cleared.
  • There were a large number of unsatisfactory and long (1-4 hour) calls to the "Help Desk". e.g. 18-Sep-2012 following Technician no-show.
    • 'A' was repeatedly shunted between departments (Accounting, Technical, ...), with no-one taking responsibility. The call finally dropped whilst 'on-hold'.
    • No evidence on subsequent calls of any knowledge of previous calls. Every call was a return to the "pass the parcel" with no person/department taking responsibility.
  • A Telstra complaint was lodged (04-Sep-2012), 'A' was given a "trouble ticket" number and told to contact Technical Support (number supplied). [[Two people assigned to the case (?), with promises to call-back within 24 hours.]]
    • Tech Support called (11-Sep-2012), on-site visit booked for following week (17-Sep).
    • Being able to speak to someone with "English as a First Language" had been an immense relief to 'A'. Finally their concerns were noted and seemed to be taken seriously.
    • Neither person called 'A' back within 24 hours.
    • When contacted, the complaints folk said they'd tried to contact 'A' using an incorrect phone number, one 'A' had never held. No apology was made for this. The complaints people could not correct the database error.
      • Having the number corrected took a good deal of time and effort in itself. Multiple departments claimed "can't do it" or "not my area".
    • 'A's mobile phone number has been registered with Telstra as their primary contact point for more than a decade. Why were any of the databases incorrect?
  • After this (mid-late Sep-2012?) a very confident Telstra employee rang and identified themselves as "Level 3 support" and embarked on a very long and trying support call. They reassured 'A' that they could and would fix the fault.
    • Under instruction, the replacement modem was connected by 'A' and failed to work.
    • When the original modem was reconnected, it failed to work as well.
      • The service was now non-operational and the support person left it that way.
      • No apology or explanation was offered.
      • The "support" person did not book a recall or ever call back.
    • 'A' was nonplussed: Telstra had oversold their competency and destroyed a usable service without progressing resolution of the fault.
  • 'A' visited a local Telstra Shop (26-Sep-2012). Wished:
    • a credit for the time the service was not provided, and
    • to cancel the cable internet service.
    • 'A' was told that because of the technician visit arranged for the next day, the service could not be cancelled. The Telstra Shop staff were not interested that the fault had not been fixed in a month.
    • 'A' had wished to speak, as a shareholder, to someone senior about costs to the business for the fault. The manager was not present, no meeting was organised.
    • 'A'  had wished to request checking and correct all related account and service records. This was not organised either.
  • 'A' purchased a Telstra prepaid wireless modem from Australia Post (26 or 27-Sep-2012), unable to get working after spending time with Call Centre. Device returned. (date?)
  • 'A' bought a Vodafone prepaid wireless modem from Australia Post (26 or 27-Sep-2012) and after a few false starts, got it working and regained their Internet service.
  • Telstra Complaints officer called next day (27-Sep-2012) to say "we're working on it".
  • Telstra sent a standard e-mail survey following up on the prepaid wireless modem (bought 27-Sep-2012).
    • 'A' detailed their disappointment in Telstra service and invited them to call.
  • (02-Oct-2012) A Melbourne based Customer Service rep.. 'R', called 'A' about the wireless modem and the on-going fault. 'R' said they would ring the next day.
  • (02-Oct-2012) The Teltra Complaints Officier assigned to 'A' called saying another person in their section would contact 'A' later that morning.
    • No call was received.
    • 'A' left messages that afternoon and the next morning. These were not returned.
  • (03-Oct-2012) 'R' rang 'A' in a conference call including a technician , 'J'  in Melbourne. 'R' had to leave the call early, with 'J' spending an hour on the phone with 'A', attempting "a manual override" of the Error Code. This required long waits and providing the hardware address of the original modem. This had to be read by 'A', 'J' did not seem to have this on record.
    • This over-ride appeared successul at the time.
    • The connection failed overnight.
    • It seems to be working today.
    • How will 'A' know the fault has been cleared?
      • They currently believe the fault is rectified.
    • Why wasn't this done on, or just after, the first site visit, six weeks earlier?
      • Why the long wait and run-around?
  • After the apparent resolution, 'A' had multiple calls from people within Telstra, all very excited the fault had been fixed.
    • None offered an apology or any compensation, some seemed to claim direct credit.
    • None offered an explanation of either the Technical fault within their systems, nor what had gone wrong with internal Telstra processes and automatic systems to cause the multiple faults suffered.
    • None offered a "magic phrase" to be repeated to Technicians and Help Desk about the fault should it recur.
    • No on-going "trouble ticket" number was given to 'A', should the fault recur.
    • No-one offered a shortcut for service if the fault recurred shortly.
  • After the overnight service disruption, it wasn't clear if one of the other Telstra personnel had undone the "manual override" with an individual attempt to rectify the fault they'd claimed.
    • There was no evidence of good co-ordination amongst the various Telstra "Silos".
  • 'A' had concluded on 3 October, that  Error Code CCP0012, is not a technical problem, nor is it an accounting problem, but an Activation error problem and simply a code that needs to be removed from the Telstra system to allow the modem to connect and activate.
  • 'A' raised a complaint with the TIO (Telecommunications Industry Ombudsman) (03-Oct-2012) sending their records of the incident.
  • 'A' had been originally told "there is construction work in your area, a cable may have been cut". This seems to have been a deliberate, misleading statement.
  • Telstra did not give any hint that after a month:
    • That the fault had been escalated
    • That for failing to provide the service, they would rebate 'A' the service charge.
    • No offer was made to supply a temporary service, such as a 3G USB modem.
  • 'A' had had to cancel a number of important business and personal meetings to wait aimlessly for a Telstra technician to attend on multiple occasions.
    • No option for an increased priority owing to the long-standing nature and difficultly of the fault was offered.
    • Telstra would never offer better than a 4-hour window for any attendance. They never scheduled 'A' at the beginning of the window, always near the end, or didn't attend.
Psychological Dimension: Induced Pain and Suffering

It's also worth noting that completely out of character, 'A' suffered extreme agitation, frustation and desperation at both the impenetrable wall of "service" and the inability to be heard, treated respectfully and to get a resolution to a service that had become necessary for conducting their business and life.

This isn't a "minor annoyance" or idiosyncratic: there is some very deep human psychology involved.

There's a branch of psychological therapy that relies on our limbic systems' (the cingulate nucleus) response to attaining goals: pre- and post-goal attainment happiness, two very distinct and important phases.

For all humans, striving and overcoming challenges is innate and core to our psychological well-being.

Consistently setting goals and achieving them isn't just "nice", but necessary, for our continued happiness and psychological well-being. Goal Attainment forms the basis of various powerful approaches addressing Depression and other conditions.

Knowingly and uncaringly forcing people into powerlessness and frustration would, in an OH&S workplace setting, be illegal: employers are required in Australia to provide a Safe Workplace. Deliberately causing employees harm, physical, emotional or psychological, is illegal and attracts civil penalties, as well as curative support for those affected.

I'm not sure if current OH&S law can be extended to customers. If so, companies like Telstra which seemingly have a policy and strategy of blocking communications and frustrating customers, would face considerable penalties...

The positive effects of Goal Attainment means the inverse, preventing people from achieving goals, is devastating, more so for high-performing individuals like 'A'. It can be categorised as "cruel and unusual" treatment, especially if intermixed with multiple events setting up false hopes and then dashing them. The human response in this case is even more profound and damaging.

Systemic Failures within Telstra

This whole adventure was unnecessary and presumably preventable: some automatic system failed when a new Cable Modem was ordered and incorrect configuration data uploaded to an operational system, without detection, audit or correction. Who has been charged with finding the root cause?

The final fix, a "manual override", should at worst, have been done the next day by the technician 'R' in Melbourne, prompted by the call from 'A'.

If Telstra's fault resolution system had worked properly, the fault would have been automatically passed to 'R's section as soon as the first technician recorded the Error Code.

  • Telstra has no fault escalation procedures, conclusively demonstrated here.
    • After any fault has been open/unresolved for two weeks, it should have been escalated to the Head of Operations for the State.
    • Any fault that is due to an internal process failure, like this, should be immediately escalated to senior officers with full cross-organisational authority and access to diagnose the root cause and initiate permanent prevention measures.
    • Own-goal process faults like these need to be reviewed, tracked and addressed by the CEO and their Senior Management Team.
      • They threaten the viability of the whole business and sufficient Responsibility and Authority only comes together at the top of all Silos, the CEO and their team.
  • High-value multi-service clients are treated worse than low-value customers: there are demonstrated errors in Service and CRM databases.
    • There is a major deficiency within Telstra: nobody is checking and correcting these service records.
      • After the "IT Transformation" project, it was known that high-value customers could not be automatically transferred.
      • To have stale data in multiple locations (service address, customer contact) says the database is seriously compromised, leading to many costly preventable errors.
    • Who is responsible and accountable for Data Quality, and do they have the Authority, will and budget to force records to be corrected?
      • This appears to be a major organisational failing and oversight.
  • Correcting faults in Telstra records is onerous and time consuming for Customers, it should be simple and easy.
    • It cannot be done in real-time whilst speaking to Service Reps, or
    • Service Reps are poorly trained or refuse to execute their tasks.
  • Telstra shareholders are treated no better than anyone else.
    • This is a marketing opportunity going begging to create engaged and supportive shareholders. Telstra has one of the largest 'Mom and Pop' share registers.
      • Service discounts, special offers and loyalty bonuses are possible.
      • Special service and access arrangements for shareholders would encourage them to give all their business and that of their immediate families to Telstra.
  • Telstra BigPond seems to offer nothing like the Telephony Customer Service Guarantee (CSG).
    • The Internet is a vital lifeline personally, professionally and in business for almost all Telstra customers now.
    • Decent Service Guarantees would match Consumer expectations and usage, as well as provide Product Differentiation.
  • Telstra's Offshore "Help Desk" with ESL speakers are counter-productive, especially for complex, long-running faults.
    • Whilst possibly tolerable for simple tasks and "script driven" data acquisition, they are a Nett Negative Value in this situation and many others. Saving money on Help Desks may be illusory and detrimental to the whole business.
      • Allow Customers to choose more expensive support options:
        • Like Airlines, offer multiple levels of pay-for-service, allowing the business to maximise profits by offering multiple price-points. (No "money left on the table").
        • Higher cost support could be automatically included as 'upgrades' for high-value customers, as Banks do.
      • After two calls on the same fault, automatically direct the Customer to a specialist Held Desk with a single person assigned and responsible for achieving Customer Satisfaction.
        • Instead of measuring "time to finish or transfer call", complex calls need to measure "Time until Customer is fully satisfied". Only the Customer can close   complex faults.
  • The Help Desk practice of "pass the parcel" is frustrating to Customers and Counter-productive as it causes significant Brand Damage.
    • Somebody within Telstra needs to be responsible for detecting, monitoring/reporting and preventing this situation.
    • Ideally, the phone system should track customers who are passed around and offer them a "circuit breaker".
  • Making Customers wait on Help Desk queues for hours serves no purpose other than weeding out those with better things to do, prompting them to look for alternate service providers.
    • Long Help Desk delays are an invitation for Customers to Choose Another Carrier, a tactic which would not impress shareholders one iota.
    • Under provisioning Help Desk service staff only serves to reinforce the stereotypical image provided by Lily Tomlin in her "We're the Phone Company" sketches. This is against the best interests of the Business.
    • This reinforces Telstra as a Toxic Brand to Customers. Whilst Customers have no better place to go, they will tolerate it. Given the choice, they will flee, never to return.
    • This is known, preventable Brand Damage at its worst.
  • Complex faults are slow and difficult to solve. This isn't simply a Customer Service and Brand Damage issue, but very expensive to the organisation.
    • This fault cost $5-10,000 more than it should have.
    • We know that this wasn't a one-off and that other Cable Internet subscribers were affected, but their faults weren't resolved.
  • The initial problem, the non-supply of an ordered service, was never addressed.
    • How much business can an organisation deliberately throw away and survive?
    • It seems nobody is directly responsible or accountable for this lost revenue.
  • Telstra, if it wants to engage and retain customers, must never internally cancel a service order without contacting the customer, explaining the situation and offering alternatives, It's an opportunity to "upsell" the client.
  • There is a major fault with the Technician ticket system: The first on-site Technician should not have been able to pass a known, unresolvable fault back to general enquiries.
    • Telstra confirmed that "TRG":
      • knew of the "Error Code CCP0012" fault,
      • that it affected multiple customers,
      • presumably they had no idea of the immediate or root cause, and
      • they had no idea of when they would be able to fix it.
    • In ITIL-speak, this was a Severity One Major Problem, but wasn't classified as such.
      • An appropriate organisational response would've been to establish a war-room comprising the State Heads of Branches and Senior Line-of-Business Managers.
      • Following the successful work-around, all other faults associated with the Problem should've been corrected.
      • Investigations initiated as to the root causes (automatic systems and processes) responsible for the error and means of detecting recurrences and costs of prevention measures.
  • Telstra's Problem Management is either deficient or non-existent.
    • Problems are not Faults, but the cause of one or more Faults.
    • To have a Known Problem not detected by the Fault Ticket Handling system is a major Professional failure that should be explicitly investigated and reviewed.
    • All the Help Desk and Technical Systems should've found an outstanding Problem with "Error Code CCP0012", with a known workaround ("manual over-ride").
      • The systems and processes need Review and correction.
      • A major internal inquiry is needed to uncover the root causes of this meta-failure.
  • Multiple Telstra employees were contacting the Client unbeknownst to one another.
    • This is the compelling reason for a CRM, a "single Client Communication Flow".
    • This definitively failed, either because the CRM was faulty, or it was bypassed or procedures ignored.
      • All of these are cause for deep concern and deserve an inquiry.
  • The inability of the Complaints Officer to progress the issue, or identify it was a Known Unresolved Problem, means their systems and/or processes are deficient or faulty.
    • This is a major problem deserving immediate attention.
  • The fault was only resolved accidentally when a Customer Survey person became involved and somehow was able to refer the fault to a diligent, competent and active Technician.
  • The unidentified "level 3" support person that caused the service to fail entirely should be found and castigated, as should the many service personnel who failed to follow-through on the fault.
    • These actions are consistent with a widespread attitude of "Care Factor: Zero", inimical to resolving faults or preventing faults through good Problem Management.
    • The one Technician who persevered should be found and commended.
  • The consistent lack of apology to the Client, the lack of any offers to provide an alternate service until the service was restored or anyone offering the statutory minimum (under the TPA/CCA) of rebating service charges show a systemic failure in even adequate, not good, Customer Service training and knowledge of legal requirements.
    • At a minimum, this is a systemic Training failure.
    • It indicates that nobody is monitoring, measuring and reporting on general levels of Customer Service and evaluating adequacy of Training.
  • The most critical and over-arching and pervasive Failure is the identification and investigation/analysis of massive cost over-runs on service faults etc:
    • This fault cost the organisation an unnecessary $5-10,000, more than the $100/year service margin could ever return.
    • This will not be an isolated occurrence, many of these will be eating away at Profits and turning away Customers, particularly high-value long-term Clients.
    • This only got resolved through an accidental interaction, the Customer Survey person who bothered to follow-up on the feedback. This bodes very poorly for the future performance on the organisation.
    • By rights, Telstra should have standing reports with automatic escalation:
      • Identify full internal cost to resolve faults and other service issues.
      • Report and escalate excessive fault resolution costs to Senior Management.
      • Mandate Root Cause Analysis  (RCA) of the "Top Ten" faults found in the RCA's.
      • Require the CEO and their Senior Management Team to track all "Top Ten" issues and regularly report to the Board on progress and problems identified.

The irony is that I shouldn't be writing this analysis at all. None of this should've happened in a well-run organisation that cared sufficiently for its Customers.

The tragedy is that Telstra will probably continue "Fat, Dumb and Happy" for the next 10-15 years in blissful ignorance of this piece and then wonder why they are "suddenly" losing Customers, disproportionately their most valuable, at an accelerating rate.

2012/10/02

The Klingon Guide to I.T. Management

My mate the DBA, whom I think writes wonderfully, coined the idea of "The Klingon Guide To Management" - not everyone might be pulling in the same direction within an organisation, not all agendas and rules may be stated and overt and those you thought were your friends may be elsewise.

I only recently came across Prof Fred Brooks latest book, "The Design of Design: Essays from a Computer Scientist" (Brooks first described The Mythical Man-Month, "adding more people to a late project only makes it later", when he wrote on the lessons he'd learned being in charge of developing OS/360 for IBM in the early 1960's). He still has useful new insights on Project Management and other Computing/I.T. topics.

Chapter 4 of The Design of Design is titled "Requirements, Sin and Contracts". He lays out nicely the human frailties (even 'sins') that make Real World Project Management much more difficult that the Ideal World assumed in the Rational Theories of Project Management.

  • Clients can be greedy, unreasonable, capricious and not pay or play fairly.
  • The Architect and Designer may have different agendas to each other and not always act in the best interests of the Client when acting as 'agents'.
  • Builders often don't have the commitment to quality, budget and schedule that the Client, Architect and Designer expect or desire.
  • "All Players are honest and truthful and communication amongst them is excellent". Or: Egos never get in the way.
My DBA mate when told this, countered with: "You know what's wrong with ITIL, don't you?"
Q: "What?"
A: .... I can't remember what he said, I was doubled over with laughter, it was so good and so true.

It was along the lines of "Everyone is competent, on the same page, helpful and cares about results".
Nope! Not within a Bulls Roar. Not seen by either of us in any Real-world organisation of more than two people.

It's a nightmare in most I.T. Ops organisations:
  • Big Ego's and on-going vicious internecine wars ("Office Politics") are the norm.
  • Finding Competent or Engaged staff is unusual, finding both in the one person is exceptional.
  • For all those of you who've rung a Help Desk, you understand "Help" has a special meaning within the I.T. Reality Distortion Field, or "It's not Help as we know it, Jim".
  • Recalcitrant Clients, Programmers and Users and Clueless Project Managers are to be expected.
  • Denial and Avoidance and  Blaming, Placating, Appeasing are the normal emotional responses of Management. The more Senior the Manager, usually the more extreme the disconnect with Reality.
  • Project Managers often get "performance bonuses" to motivate them in achieving features, budget and schedule. What you get instead is bullying, intimidation, threats and lies directed at staff and vendors and "snow jobs" for those up the chain. Getting those "bonuses" take precedence over all other Stakeholder Requirements... Which doesn't improve the result for the Client or Organisation.
The People Side of I.T. Ops and Projects overwhelms the Technical. The only saving grace is that I.T. people are usually very poor at Office Politics, so in spite of them, things occasionally happen.

There is a real need for The Klingon Guide to Management, especially in I.T.
I'll keep my fingers crossed for it to be written.

"My enemy's enemy is my friend". Nope! They're both your enemy, destroy them both with all means available! Ahhh, if only I'd known that when a youngster.

An addendum: Another good friend volunteered two things about problems in IT Ops:

  • Are they competent, diligent, helpful and prepared to listen/debate (vs arrogance)?
  • Do they recognise and understand the problem? Are their responses considered and supported?
    • Not acknowledging problems, being defensive, blocking or deflecting ("we didn't change anything. What did you do?") are classic responses we've seen in IT Ops when asking others to repair services under their control.
    • Is the solution or facility they offer or want backed by need or evidence? Very often what gets done comes down to a battle of wills. Confidently asserting you position is what makes you right, not facts and evidence. Evidence based repair and remediation is the exception, not rule.
      • One client I advised chose to ignore my written report and purchase a $1MM "special" package from a vendor - a less than 50% List Price "End of Fin Year" deal. Salesmen do deals to make their quotas, not meet client needs... My recommendation was for a $300,000 system, which they bought within a year for another project.


2012/09/24

NBN: Coalition cannot cost its proposal

Malcolm Turnbull has finally addressed just what he meant by "sooner, cheaper" in an ABC radio interview and an article based on the interview. [Update: 7:30 Report Interview, Coalition B'band Survey]
TONY EASTLEY: The Federal Opposition says it won't be able to provide a fully costed broadband policy by the next election but its plan will be cheaper and completed sooner that the Government's National Broadband Network.

The Opposition's communications spokesman Malcolm Turnbull says information from a survey to be launched today will help a future Coalition government decide which areas to prioritise for faster broadband services.

NAOMI WOODLEY: How long will it take for the Coalition network to be rolled out? NBN Co's timeframe is around 10 years. Yours is sooner, can you say how sooner?

MALCOLM TURNBULL: Well it will be a lot sooner. But let me just first say that NBN Co says its target is 10 years. There are many people in the industry very close to the NBN who believe it is more likely to take 20 years.

The approach that we will take in most of the built-up areas of what's called fibre to the cabinet or fibre to the node, the experience around the world is that takes around a third of the time of fibre to the premises, sometimes even less.

2012/09/14

NBN: Coalition supports Free Markets not Monopolies

Mr Turnbull wrote a commentary on a speech by the Chairman of NBN Co, "Policy Choices in Project NBN".

Mr Turnbull engages in the usual political rhetoric, specious comments and gratuitous insult, summarised at the end, but does raise two points that need addressing:
  • Monopolies, and
  • Provision of new services during the transition from Telstra to NBN Co., "greenfield" installs.
In a single paragraph, reformatted here, Mr Turnbull makes what I see as a substantial, multi-layered policy statement.
The key words here are “without making the subsidies explicit and transparent.”
And therein lies a very big difference in philosophy.
Not only do we prefer competition to government monopoly, but we believe that subsidies should be explicit and transparent.
We are thoroughly committed to providing access to broadband to regional and remote Australia at city prices –
   so there is no discrimination by reason of geography –
but the cost of doing so should be transparent.
(precedent) What, after all, is the USO?

2012/09/12

NBN: Coalition support for Free Markets vs Monopolies

Mr Turnbull wrote a recent commentary on a speech by the Chairman of NBN Co, entitled "Policy Choices in Project NBN".

[Long. 2850 words total, 2000 words of mine]

I made this comment on his blog:
Malcolm,
You keep on top of the debate and are continually moving it forward. Excellent work as both a Shadow Minister and prospective Minister.

I’m impressed with your criticisms of the NBN Co rollout. Informed and useful.

Here is not the place to nitpick the Coalition policy. You have one, have made promises that can be evaluated and have publicly committed to continuing a National Broadband Network.

Thanks for being so active in this debate and not allowing the incumbents to be complacent or go unchallenged. Well done.

regards
steve jenkin
Here are my criticisms of his blogpost.

2012/09/10

Future of Big Iron servers and Expensive Databases

This article came across a list, what follows was my response.
IBM and Oracle Present Rival Chips for 'big Iron' Servers
Wonderful to see the Dinosaurs still dukeing it out.

The eDRAM from IBM for the L3 cache is a big move. Like when they figured out how to use Copper in chips and reduced power use, hence heat production significantly.

I suspect we're seeing a replay of the late 1980's demise of mainframes... Not fast, not universal, not complete, but 90+% of the business goes away, killing weak supporting businesses.

Everyone but IBM's System Z and Unisys ClearPath went away - or into emulation.
[Clearpath = emulation on Xeon of 2200 & B-series]

In my view, two forces have converged to push these high-end niche processors into irrelevance:
  • Patterson's Brick Wall (2006):
    • Power Wall + Memory Wall + ILP Wall = Brick Wall
  • "infinite" IO/Sec and virtual-RAM with PCI-SSD. eg Fusion-IO
With cheap PCI-SSD by the Terabyte, the majority of Apps/enterprises don't need:
  • Big Iron Databases
  • Big Iron Storage Arrays and supporting SAN's
  • Big Iron multi-chip fast uniprocessing cores.
A lot of the complexity of Big Iron DB's (like Oracle) is aimed at achieving "speed" in the face of low-performing HDD's... [slow IO/sec, not streaming throughput]

If the whole of a relational DB (tables) fits in memory (or fast Virtual memory), then doesn't the DB become very simple, modulo ACID tests and writing "commits" to persistent, high-reliability storage?

Which means we might start seeing a bunch of in-memory DB's, like NoSQL, but for normal-sized DB's (1-5Gb), not large collections.

There's an economic rule on product substitution that led to the relatively quick decline in IBM mainframe sales:
  • when the capital expenditure on a substitute is less than the operational costs of the current system, barriers to adoption are removed.
    • capital costs are 'sunk' and can't be recovered.
    • To realise savings, you have to wait for the next upgrade or refresh cycle.
    • But the conversion costs have to be factored in, and incumbent vendors take care to price upgrades, even "forklift upgrades" (complete replacement) under the total cost of moving to a new solution. [Used to advantage by Tier 1 Storage Array vendors currently].
    • But Operational Costs, like maintenance charges, aren't 'sunk'.
      • They are due every year.
      • When a whole system is less than recurrent costs, businesses can quickly and easily justify the change.
      • They write-off the CapEx for the old equipment and wheel it out the door.
      • Usually "Big Iron" hardware has zero residual value, even when 2 years old.
      • When the equipment is on the cusp of being obsolescent, it is worse.
      • Or they have to figure out how to break the lease.
      • Unless they went into debt to fund the CapEx, they can move away quickly.
Hardware maintenance fees are typically 15-20% of capital costs.
Whilst Oracle licensing costs are beyond me (I don't track them) - but are becoming a major component of Enterprise Computing costs.

How many "little" DB applications need to succeed with two low-cost ($10k) servers, 3 SATA drives each in simple RAID and 1 Fusion-IO board, run as H/A with an in-memory DB?

If organisations can build a complete, high-performance, high-availability, simple-admin solution for $20-$30k per group of DB's, they can afford to deploy them immediately based on direct maintenance savings.


Follow-up comment:
Intel are looking over their shoulder at becoming dinosaurs. Maybe I won't live to see it, but ARM servers could very well do in the x86_64.
Ed
And so they should for most general purpose computing.

I remember seeing John Mashey of MIPS talk in 1988 where he plotted CPU speed for each of ECL, bipolar and CMOS technologies. ECL had been overtaken by then, bipolar was due to lose the lead within a few years.
The 486 in 1991 was a complete system-on-a-chip and changed the landscape.

It answers the question "Where did all the supercomputers go?"
A: Inside Intel. [and Power and SPARC. possibly Z series]

The Intel chips seek "maximum performance" - they pull all the tricks that super-computer designs used, and its is that technology that is approaching Pattersons' "Brick Wall" [heat, memory, ILP]

And as an aside, GPU's are filling the "vector processor" niche of CDC and Cray.

ARM has pursued a very different strategy, more based around 'efficiency': MIPS/Watt

So, while I agree with you, I think the situation is nuanced.

ARM processors are obvious choices for low-power and mobile/battery devices.
Because of design simplicity (small PSU, no CPU-fan) and smaller size, they'll become more interesting for low-end PC's, especially portable devices.

There is a company, Calxeda, now producing high-density ARM boards for servers.
They are hoping to leverage MIPS/Watt for highly-parallelisable loads, like web-servers.

But I can't see anyone taking on Intel soon in the supercomputer-on-a-chip market.
It's not just servers, especially for large DB's, but workstations and 'performance' laptops.

The problem with that evolution of the market for Intel is ARM taking sales from multiple market segments. Seeing that Winders-8 will run on ARM, we might see the end of WinTel for low-end & mid-tier laptops.

As a company, can Intel survive such a radical change in demand for its major product line?
Will its work on MLC flash fill the financial void?

I've no idea how that will go.
But like you said, ARM is going to shake up even the Intel server market.

The "secret sauce" that the ARM architecture has is that it's a licensed design.
Although chip design companies might not own or be able to access chip FABs within 2 or 3 design cycles of Intel, they can produce highly optimised and use-case targeted chips.

Which Intel can't do. They are focussed on the bleeding edge of CPU performance and FAB design.

Manufacturers like Apple/A5 and Calxeda can produced ARM-based designs that can outperform Intel-based systems by an order-of-magnitude on non-MIPs metrics.

As Apple has shown, there are very big markets where raw MIPs isn't the "figure of merit" in designs.

2012/09/05

Recruiting FAIL: Part 3. ITCRA complaint


Lodged an formal complaint with ITCRA [IT Contract and Recruitment Association], the Industry Body for Recruiting companies.

Several of the Agents actions were serious breaches of the ITCRA Code of Conduct.

In 30 years of dealing with Agents, this guy is by a long margin, the worst I've come across. Not just incompetent  or "poor with details" such as misspelling Workseeker name on contract. Also demanded an apology as pointing this out was "offensive".
Background:
This was a *permanent* position for a senior IT staffer.
Workseeker required to move interstate.
Agency approached workseeker, no job application ever made by workseeker.
Job never advertised with a start date, no immediacy ever stated. Comment "expected to take 2 months to fill position" [by October]

Complaints.

Misleading statements:
1. Confused an email from himself to workseeker with a written offer from the Client. Insisted this was a "Job Offer", implying a binding contract.

2. Promised help finding accommodation for relocating interstate, none provided. Having local accommodation was always a condition of accepting the position.

Harassment or cyber-stalking:
3. One incident of a dozen calls/texts in 30mins.
Asked, in an email, to desist.
Repeated again later with more very inappropriate and abusive statements.

Privacy Act breach.
4. Used Referee's contact details for purposes data not supplied for. When workseeker would not answer Agents calls, Agent rang referee to complain and abuse.

Possibly Criminal over-stepping
5. When workseeker had failed to receive a contract 10-days from 1st proposed start date, informed Agent of inability to start due to personal circumstances.
Agent then harassed and browbeat client.
Seriously overstepped by continuing to demand exactly what the personal circumstances were.
Justified by saying "I need to explain to the Client".

This was uncalled for as the Client had no urgency on filling position, nor was there any advertising start date.

6. General Incompetence and lack of attention to legal details
a. Misspelled Workseeker name on contract. Demanded a  written apology to Client for pointing this out.
b. Sent 7 emails in 15 minute period informing workseeker of Job Offer. Kept issuing 'recall' emails for incorrect offers sent.
c. Commented that Agents' manager was continually chasing him to comply with company requirements and document all communications.

7. Failing to act on Instructions.
Workseeker requested 3 times in email that Agent ask Client if 9-day fortnight [with RDO] would be acceptable during first 3 months while transitioning to Sydney.
Agent never asked Client [confirmed].
Client was happy to comply when asked.
Agent deliberately misled workseeker by stating the job "was full-time only" and RDO's were not possible.

Emails/SMS's documented at:
[http://stevej-on-it.blogspot.com.au/2012/09/recruiting-fail-how-to-foul-up-employee.html]

2012/09/02

Recruiting FAIL: Part2 - Red Flags and Lessons Learned

This experience was unpleasant enough that I took down my LinkedIn account with around 300 contacts, and resolved I wouldn't look for work as a System Admin again.

Whatever I do in looking for work, it's wrong, there is no point in pursuing a strategy that's failed me time and again for over 10 years.

I took a clear decision to provoke a crisis by sending my 'problem' email.
I had accommodation organised and paid for the first week, had packed and organised myself to start on the Monday and a plan, if a little shaky, to continue.

I've learnt a harsh lesson, which means "better to abandon earlier than later":
Things go on as they start. or It will only get worse, not better.

Lesson 1: If it's important to you, get it in writing early on. More so for "dealbreakers".

I never got Slippery Sam the Agent to make a written commitment on what he was promising to deliver. He, and the company, couldn't be held to it.

Lesson 2: Relocating cities is a Big Deal. Allow time, Plan the Move and organise Reconnaissance.

Driving a few hours up the road for an interview isn't the same as moving your life. Even if single, you have to devote a decent chunk of effort to the task. It will take time to do properly.

Lesson 3: You can't make these decisions alone. Talk them over with a friend.

If I'd talked through my decisions and the way I was being treated with a friend, I may have slowed the process down and set a much better process and plan to move in place.

Lesson 4: Be wary when there's a Big Rush and you're not asked when to start, but told.

There was never any hiring date from the Agent or company. They seemed to turn it into a huge rush, but hadn't declared there was any problem that needed someone there Right Now!

There wasn't a contract negotiation, there should be at least a start-date negotiation or specified in original request.

Lesson 5: One Red Flag is enough. Two is a dealbreaker.

Agents and Salesmen will always come across as Great Friends. Which you are until they got your signature, then you're a pariah.

The first Red Flag should've been the apparent haste (Job Spec on Wed 8-Aug, Interview on Mon 13-Aug, organised on Fri 10-Aug).

The next, the lack of a specified start date.

The Red Flag, par excellence, was rescheduling my interview a) earlier and b) on the day.

Slippery Sam the Agent went on to harassment (12 calls/texts in 30mins), browbeat me and completely overstepped the boundaries by demanding I explain my personal circumstances.

That little escapade was, in retrospect, an Instant FAIL.

Agents are there to facilitate the engagement, not beat-up on you and heap on abuse.

Recruiting FAIL: Part1 - How to foul-up employee engagement.

It began, as I recall, on a sunny winters afternoon in August, a Tuesday.
22 days later it had ended acrimoniously with an SMS and my email in response, and the removal of my LinkedIn account to avoid such agents/events in the future.
Please pick up your phone and talk with me steve. Like adults, let's discuss this. At the moment, you are really damaging my relationship with my client which is not fair and not right. I have really done all I can to help you and you won't even talk to me. SMS: +6145206812. 29/08/2012 15:42:49
and
To: B and C
Subject: Harassment
Date: Wed, 29 Aug 2012 16:43:44 +1000

B,

it's bad enough that against my express wishes you've been bombarding me with calls and texts - that amounts to harassment.

BUT TO CALL MY FRIEND??? What the HELL were you thinking??

Yes I'm ignoring your calls, not because I'm petulant or sulking, but because:

a) I've been doing stuff today, including taking a load of stuff to the tip (loading, driving, unloading, driving - not available to talk), and

b) because there is only ONE thing of interest you can say to me...

Which is: "This is how we can work this out..."

Unless you've got a plan to get me reasonable temporary accommodation while I find a place I can sign a lease on, then there's nothing new to be said.
I can't afford to commit to $10,000 or more in a lease on the chance I'll still have a job in 6 months.

I'm happy to work at C/O, I like C and his team, I think the work would be interesting and think I could make a positive contribution there.

There is just one thing, no more, standing between me and starting there and that's I don't have a place to stay and you haven't helped.

If I had applied for a job in Sydney via you, then it would be my problem to look after myself, pure and simple.

But that didn't happen.

YOU approached ME.

If YOU want me THERE, YOU have to make it happen.
Moving states, you of all people should know its a big deal.

So far you've made empty promises and hung me out to dry...

So - STOP HARASSING ME.

The only message I want to get from you is "It's fixed..."

Otherwise, we've said all that needs to be said.

steve
I precipitated this crisis with an email just before Noon to the Agent (B), the manager (C) and (F) the person in CO that'd sent me the contract. B was very upset that I'd sprayed a message all over CO - told by my friend G, whom he'd first contacted as a referee and then again after this note when I declined his repeated calls.
To: B, C and F
Subject: problem
Date: Wed, 29 Aug 2012 11:49:41 +1000

I'd like to start by reminding you that I didn't apply for a job with CO - that you've chased me, being insistent to the point of browbeating.

You've known from before you talked to me that I had to move cities to take up your position.

I made two things very, very clear at the outset:

- I do NOT have any temporary accommodation in Sydney, not family or friends I can crash with for a week or two, and

- I required assistance to find a place. From out of town, it's very hard to find anything, especially if a constrained timetable is imposed, as you've done.

Rephrasing this: My employment with CO has always been conditional on finding accommodation, and you've been aware of this.

Two weeks ago, I was promised introductions to estate agents and assured it'd present no problem finding me somewhere to live, presumably reasonably before your selected start date.

As yet, that promise hasn't been fulfilled.
With 2 and a bit days to find a place, I really don't think I could find a place, and certainly not something that I could afford or that I would want to live in.

You've run out the clock on me. I've no idea why.

This week, there was a news item about a murder in a boarding house and I thought: "never again".

A number of times when I was contracting and working in Sydney, I stayed in these unpleasant little dives. Since the news item, I won't be forced for no good reason, to go there again.

I've not been impressed with either how you've generally communicated with me, nor the seeming absence of internal communications at your end.

So I need to spell this out:
  • You need to get back to me on this.
  • You need to come back with an explicit proposal, either of accommodation or something that's guaranteed to lead to it.
  • I expect ONE professional communication back to me on this.
    • no browbeating or harassment, this is a problem of your devising (I got 12 calls/SMS in 30mins when I walked to the Post Office)
In normal engagement scenarios, asking people move inter-city involves a bunch of things you haven't done:
  • flying the candidate to the in-person interview
  • making allowance for the relocation in the start-of-work date
  • paying or, or contributing to relocation expenses, every airfares
  • providing temporary accommodation, usually some months, while the person finds new accommodation
  • providing time-off to search for accommodation and settle affairs, as needed.

regards
steve jenkin
The manager, C, responded with this, making it very plain he wouldn't help me, personally or corporately, nor attempt to negotiate a solution.

It fits my saw "your priorities are what you do, not say".
Subject: Re Harassment
Date: Wed, 29 Aug 2012 16:51:44 +1000

Hi Steve,

We CO have no way of providing you with a solution for your accommodation here in Sydney and as I mentioned in our phone conversation before, although it is in Agency's best interest to match CO with the right candidate for our job posting, it is not their job either. Unfortunately we will have to miss out on this opportunity to work together.

I wish you the best of luck in your future endeavours.

Cheers,
C
There had been a fuss previously with the contract. I'd been invited by F to ask questions.
The response to my questions, I found fantastic, as in beyond belief.
  • I was told both B and F found it "offensive" and condescending, but was never told just what I'd written they been offended by.
    • The CEO was shown my email (itself an interesting move, not entirely legal) and laughed. Said he liked my directness and saw no cause for offence.
  • B blew me up for writing to "his client" without permission and instructed me to never contact F directly again.
    • I eventually got B to understand that he'd not mentioned this rules earlier that day when we talked and inventing rules after the event and then chastising me for not following them was impossible logic.
    • I never got an apology for this abuse.
  • B demanded I write F an apology the next day, with him vetting it first.
    • I did so by 09:32.
    • B insisted I was to not write that he'd been involved in asking for, or vetting, my apology.
    • I objected against this, as it is a deliberate fabrication.
My questions/comments on the contract:
To: B and F
Subject: Re: Employment Agreement
Date: Tue, 21 Aug 2012 17:31:40 +1000

F,

Hate to do this to you, but my surname is singular, not plural:
JENKIN, no 's'. [At top and on signature line]

It happens a lot, which is why I emphasised it originally, I can't think how that may have been overlooked.

The contract needs to be redone to correct this.

I will sign and post a copy to you tonight (express post, delivery tomorrow) with the incorrect name noted. On my start day we can sign the new version.

My formal name for legal docs is "William Stephen Jenkin", but I'm called "Steve". No need to put one of my Christian names in brackets.

Questions:

0. Whom do I report to whilst my erstwhile manager, C, is away on leave?

1. Clause #3, probationary period.
The Fair Work Act sets a 6 month minimum for lodging dismal cases.
Would that be a more effective period than 3 months?

2. Clause #5, working hours.
The wording suggests fixed start/end times resulting in an 8 hour work-days or 40-hour weeks, yet ordinary hours are 2 hours less. Could you please clarify this for me.

3. Clause #5, flexible hours.
As I'll be relocating from Canberra, I was hoping for the first six months of my employment that I could work 9-day fortnights, i.e. 76 ordinary hours in 9 days with alternate Mondays off (RDO style). [An extra 51 minutes/day]

Will that be possible?

4. Clause #5 and #6. On-call duty and remuneration.
[snip]

5. Clause #10, Intellectual Property.
[snip]
Will the company advise me of all everything it does.
I can't see how in my position I could adequately fulfil that obligation.

6. Clause #10, Intellectual Property.
To be clear, the contract makes no mention of contributing to "Open Source Software". My understanding is that such work falls outside the scope of this clause.

regards
steve jenkin
The issue of RDO's was critical.
In 3 separate emails, I'd asked B if CO would consider letting me work 9-day fortnights for the first 3 months or so, so I might attend to business in Canberra. While I said it wasn't a "drop-dead", it was important enough for me to keep raising the issue.

B was very adamant that "this was a full-time job, CO won't let you work a 9-day fortnight".

Only I hadn't told him I'd guessed C's email address and had a conversation about this and a few other topics. C was more than happy to arrange flexible hours and an RDO. He didn't want "stress".

C was delegated to reply to my questions, which while apparently honest, didn't strike me as being well thought through, nor complete [whole email not included]:
2. It's been a long time since I've done 40 hours of work at CO. We regularly do more than 40 hours of work on a week. I am hoping that with you joining the team this will get better, but this is not a place where we work 9 to 5. Working hours are in general flexible with some restrictions given that as a team we have to ensure our 8am to 6pm support line has someone here to answer the phone.

3. We can manage that internally and informally. [RDO's]
The really important Red Flag occurred on the Friday after the Interview of Monday 13th.
The first Red Flag was C rescheduling my appointment 2 hours earlier. I was rung by B while I was driving up. It cost me $45 in parking, but I was able to make their changed timetable.

At the end, C raced out without properly terminating the interview as he'd run overtime discussing his technical problem and had a taxi waiting.
Subject: Re: FW: Confirmation of Offer
Date:Fri, 17 Aug 2012 12:00:27 +1000

B,

I'm happy to accept the offer generally BUT my personal circumstances have changed slightly since the informal offer on Tuesday/Wednesday and I can no longer start on the 27th.

With C now going on holidays, it wouldn't be operationally effective for me to start on the originally advertised date, Mon 3-Sept-2012.

I'm happy to negotiate a start date when or after C returns from holidays, but realise that CO may wish to rescind this offer to me and go with a "Plan B".

Hope to hear from you soon.

regards
steve jenkin
The first Red Flag was a total of six emails from 09:43 to 10:05 on Fri 17, first confirming the job offer, then "recalling" the email as he'd made an error, and again and again... Geting the contract details right is a necessary competency for an agent.

At 16.05, I finally received a contract from F at CO, whom I'd not heard of or from before. Mess with names and questions is above.

The Red Flag extraordinaire was a dozen missed calls and texts from B in 30 minutes while I walked down to the Post Office, sans phone, to put a signed copy of the contract in the mail (express post) so it'd be there in their office 09:00 next working day. Canberra is treated as Regional NSW by Australia Post with normal mail normally taking 2 days going via Wollongong.

I then got The Third Degree from B, demanding four or five times to know just how my personal circumstances had changed as "he had to sell it to the client". I felt abused, browbeaten and upset after that little tirade.

All of which was pretty surprising as they'd never suggested a start date, never suggested there was any hurry to fill the position and then I found out C, the manager, was going to be away just after I started.

Just to show that B did know that Accommodation was Drop Dead for me, I include an email from him.
I got another call from B around 3PM the next day, Tuesday, while I was driving back from Bowral where B said he was having a little trouble getting anything for me from an a Real Estate agent.
I specifically asked that he sent me the contact details of at least one agent.

Nothing came that evening and I waited until around Noon the next day before sending my "problem" email.
Subject: Are you free to chat today?
Date: Mon, 27 Aug 2012 14:39:30 +1000

Will you be free to chat today?

I have some people that can talk to you regarding accommodation and keen to catch up.

Thanks
B

One of the things I found "Not Quite Right" early on was a request that I take a look at a Performance Problem on one of their production systems as part of the Interview process. Turns out it wasn't fake data or a test, but I was in front of a console of a running production system, having someone else type my commands, trying to diagnose a live problem for them. Without any background, system maps or application briefing...

This had been my initial queasiness at the request.
Subject: Re: Interview Request
Date: Sat, 11 Aug 2012 19:14:18 +1000

B,

On reflection, I'm thinking there's a mismatch between this part of the Interview process (snippet below) and the job description as given...

Still going to do the interview, but I'd like to flag my concerns with you first.

The questions they are asking in the Interview are specifically:
- Capacity Planning (circuit planning, load forecasting, provisioning) and
- Performance Analysis (scalability, bottlenecks, rearchitecting, design, interactions, ...)

The closest thing in the job description is:

"A multi-skilled technical profile is required with proven ability to monitor and troubleshoot database and web server performance."

Which includes fault analysis but NOT architecture redesign and Performance Analysis.

If they're wanting a free consult (with current data), I'm not happy with that. I've done enough consulting gigs to be wary of customers asking "just one little question", then getting stiffed on the purported contract

I'd be very happy to look at data from a year ago and analyse that, then compare my analysis with what they did and what happened.
That's completely fair to both them and me.
They don't give away their secrets and I can't be conned into a freebie.

I'm not inclined to give them high-quality consulting advice for free.

Remember that I have bailed multiple large, high-profile Govt. sites out of extreme situations before, so they might just be casting around for that.

Extremely happy to be engaged by them as a consultant for an appropriate daily fee and look at any problems they give me, if that's what they truly want

Anyway, I hope they are being straight with me or just a little naive.

That's the position I'll take until I confirm otherwise.

cheers
steve
I had paid $350 to a boarding house for a weeks' accommodation. I had organised accommodation for the first week, but decided I was going to be miserable living there in a state of permanent anxiety both for my safety and if I'd have to find a new room at short notice.

Part of the problem was the uncertainty of employment, and since 2002, I've had 3 failed attempts to make it through to permanent employment. I've become very wary of Hidden Agendas of employers and managers...

Very kindly, the boarding house allowed me a 2/3 refund, as they'd not informed me of the 14-day cancellation policy. In the end, I got through this sorry mess for under $250.

In Part 2 I will try to extract the Red Flags and my Lessons Learned...

2012/08/25

NBN: Politics trumps Economics, Business, Technical and Social needs.

In a previous piece, I commented on a set of posts on the Business Spectator site, centred around Alan Kohler and Malcolm Turnbull 'debating' the Liberal plan for a National Broadband Network (NBN).

Politics is The Art of The Possible, where Perception is Everything.
The NBN is first and foremost in the Political realm, not in the Business, Economics, Social or Technical.

2012/08/23

NBN: The Turnbull Guarantee we'll never see...

Technology Spectator has written some recent pieces on Turnbull's NBN proposal I've commented on (below):
If Mr Turnbull owned the NBN he's proposing, would he offer an effective bandwidth guarantee (or your money back)? I wonder...

As far as I can see, he's proposing VDSL2 with as few Cabinets and as little backhaul as he can get away with.

Saves $$$
But, if that's his solution, in busyhour, most people won't see 5Mbps.

Will Mr Turnbull offer a meaningful guarantee of minimum performance?
Could that have

Because NBN Co is not notionally a Govt service, but a commercial Co, the ACCC might have something to say about election promises, guarantees and misleading or deceptive comments.

So, will Mr Turnbull and the Coalition offer customers a performance guarantee for the their NBN?
If they do, will he/they be liable for commercial statement under the Trade Practice Act/Competition and Consumer Act?

That would make Political history in a way nobody would want to see.

2012/08/12

Battling the WhizKids: How I kept my website up in spite of the Prime Software Contractor

1999, the lead up to Y2K was a horrendous year for me. Too many long weeks, too much travel, tight deadlines, multiple competing projects, not enough sleep and a level of exhaustion that compromised my short-term memory and took around 5 years to mostly recover from...

So at the end of 1999, when a long term friend, Peter, offered me a part-time contract looking after a few web-servers, it sounded like a doodle. Easy work, reasonable hours and good people to work with. What's not to like?!?!

I got more 80+ hour weeks, a horror project and an exceedingly difficult Software supplier to work with. But because of my efforts, the site stayed up, remained usable and our area was the only part of the overall project to not make the papers (in a bad way). Apparently the project won a Gold Government Technology Award as well.

I presented a paper on the Performance aspects at CMG-A and a presentation.

All this and a very high-profile website: the first large-scale on-line Transaction of the Australian Government: the initial ABN on-line registrations.

While the website and project were done by the Tax Office (ATO), Business Entry Point (BEP) whom contracted me ran a website hosting service out of the Secure Gateway (SGE) facility in the Edmund Barton Building (EBB), home then of Department of Primary Industry and Energy. 15-25 Agencies used BEP for their websites, maybe more. The ATO had decided when the GST legislation had been passed that it wouldn't be able to build an ABN registration website (integrated and secure) in the time-frame available and opted for using BEP.

I've never tracked down the cost of the project: BEP had bought around $1M of SUN kit for the ATO servers and firewalls. DSD security had required separate machines for the webserver and Database, with two firewalls, one to the Wild Wild Web and another between the web-servers and DB.
There were another 6-10 small SUN servers at Wizard used for testing.
And multiple Oracle and Netscape licenses.

We added an Alteon load-balancer ($10-20k?) between the firewall and webservers around Jan 2000. One of the best decisions we made.

The whole ATO project, with BEP systems could've been $2-3MM, with perhaps $1MM going to Wizard for its software.
Which, as was pointed out to me, at $5 per completed registration, was very good value for the Australian Government.

BEP was under the "Small Business Branch" of Department of Employment, Work Relations and Small Business. The main offices were in town with 30-50 people, IIRC. It wasn't long before I started working first 40 hour weeks, then much longer - banished by myself to the EBB basement, home of the SGE.

The GST was to be introduced 1-Jul-2000, and all entities wishing to pay or claim back input costs need Australian Business Numbers (ABN's) by then. ABN applications would be processed within a month, giving a deadline of 31-May-2000 to register in time. The ATO expected about 1 million (1M) applications, there were 3.3M. We processed 600,000 complete applications on-line. There were other processing streams via the Accountants electronic submission systems and on paper. The ATO project manager was to receive a bonus if 20,000 (I think) registrations were done via the Web. Their projections were for 2% of applications on-line, not the 20% at the end. After the deadline, around 50% of new ABN applications came through the website.

The irony, and I didn't spot this, was that the ATO promised two week turnaround for electronically lodged applications - via the Web or Accountants. We had another mini-rush around 15-May.

Just to make things interesting, we had to handle 1/1/2000 early on and "someone" had decided that although the systems were all patched and OK, we had to take them off-line for the rollover. I negotiated that back to "leave up a simple static page telling the world what we were doing".

And yet another wrinkle was the Federal Government Outsourcing program: they decided to sell the SGE in Jan 2000. Right in the middle of the lead-up to our deadline. All the SGE staff were let go, but not before training the staff of the successful tenderer. Who of course promptly moved those staff on.

Chaotic doesn't do justice to the initial state of the Secure Gateway under the new owners... They had their own share of political in-fighting, with one of the co-owners/directors eventually being sacked and banned from the site.

The reason I'd been hired is that the long-term Systems Administrator of the Prime Software Contractor, Wizard Information Systems, had quit. This was to be a recurring theme, people we needed left or were reassigned and unavailable. We had multiple people titled "Project Manager" to deal with. None of them overwhelmed us with confidence, good customer service or helpfulness.

The head of BEP at the time was Dr Guy Verney who navigated this who storm exceedingly well. I had my own troubles dealing with the "WhizKids", but was unaware (all) others in BEP similarly had problems. After the ABN project, the BEP staff presented their head with a "Show Cause" letter they wanted given to Wizard - exceptional within the Public Service. Dr Verney met with the head of Wizard, Tony Robey, and the letter was discussed but not served. The BEP staff were promised "it'll all be different'. Unsurprisingly, nothing changed...

In October, 2007, Wizard Information Services were put into Administration. A very sad loss to the ICT Industry in Canberra. I'd been employed by Dave Schwartz early on when he was "Wizard and Liveware". One of the best, most honest agents I'd ever met or had.

From what I'd been told, the ATO project started with a closed tender to three suppliers, with Wizard selected. The ATO project manager volunteered to me at one stage that he'd asked Wizard to requote to write the system in 'C'. Wizard had elected to implement the system in a scripting language they'd written, then charged us run-time licenses for on each platform (web and database servers).

Wizard refused to reimplement their code in 'C' although there appeared quiet adequate time and resources available.

The Application was wholly CGI (common gateway interface - start a new program for each webpage), which burnt considerable CPU resource for us on the web server. It also posed performance problems for us because SPARC's at the time could create maybe 100 new programs a second. Even with our 8-way web-server.

Half our system load was just starting the CGI processes for each page.
A simple optimisation by Wizard, to reuse a small number of existing processes, would've improved our performance considerably. But that wasn't to be.

Wizard had not thought to instrument their Application in any way or build any tools, let alone real-time, to monitor the systems or load. For a critical new application, this seemed an incredible oversight to me. Because we used https/SSL connections, the Application could have had start/stop timers added trivially, but that was not to be.

I had to invent a method to estimate the response time of the Application.
Luckily, the poor Software Engineering methods of Wizards actually helped me.
We had Unix "process accounting" turned on for security and audit reasons, so I was able to extract some reasonably detailed and fine-grained information on process execution.

The Application consisted of a first web form: "New Application or Continuation".
For 'Continuing' Applications, clients entered a number plus a password. For new Applications, they were issued a number and were requested to enter a password.

Following the initial form, clients would fill in 20-30 additional forms, driven by their individual situation, with each form containing roughly 20 fields. Each Application consisted of ~400 user-entered fields and around 100 extra system fields, such as when the Application was started (and hence when it would be automatically cancelled) and time stamps for a series of events. e.g.: completed, scheduled for transmission, transmitted, response received.

The bizarre thing that Wizard did, that made it trivial for me to identify all processes of their Application, was they handled ALL forms in the ONE program. At the start was a massive "if then else" statement or set of "go tos" to figure out "where are we up to" and "what do we need to do this time". It may have been their wonderful scripting language didn't support libraries or included common code, or they really were just that bad at Software Engineering.

Normal good design would create multiple executables, each one tailored for a single or small set of related input forms. This reduces the "footprint" of each running process, separates the various parts of the source code, allowing changes to be tested and deployed quickly, easily and with high-confidence.

So I just picked out the process accounting records that matched and made the assumption that the wall-clock execution time would be related to the "response time" experienced in the Users Web Browsers. This isn't quite right, but without extensive javascript such as Google uses, its a good first approximation.

I then added this data from the process accounting logfile scanned every minute to a free real-time graphing tool, RRDTOOL. (Round Robin Database).
I'd taken it on myself to create a simple system/health monitoring system that graphed important system load variable - including User Response time, as derived above.

This simple set of pages was one of the biggest learnings I had from BEP/ABN registrations.
On Feb-29, the ATO advertised the "business.gov.au" website, and the next day we were flooded.
Which, not unpredictably, caused the Wizard code to fail and overload our systems.

"Crashed" is too strong a word. We had an "avalanche" or "cascade failure".
Users would be entering data and getting responses in 3-5 seconds.
And then the site would stop responding - so automatically, Users would click on the form 'submit' button again after around 5 seconds. Probably multiple times after an extended delay.

The problem with that is we couldn't stop each of  those 'submits' from being processed afresh because of the combined effects of SSL and CGI. The web server queued the HTTP requests from each user on their dedicated link and processed them in oder. Regardless of what was in progress.

So just when the system was slowing down and struggling with the load, users suddenly starting madly clicking their mice... I never did track down what the Wizard code did under these circumstances. Given how they handled other "unexpected" situations, I guess very badly.

The overload was resolved in a few days by bringing forward an update of the Wizard scripting engine that improved performance around 10-fold. Which makes you wonder why they didn't do that in the first place. Wizard went on to sell systems with the scripting engine/language that had been paid-for by the ABN project to some other Agencies, notably the ACT Government. They still called the CGI script "bep.....cgi".

This event gained the Ops team three things:
  • Within 2 hours we had a purchase order for SUN for a $100,000 upgrade to the servers. We'd been trying to justify this for some time.
  • I'd expected an overload situation and written a specification for a "Busy Tone" system: when the system was flat-out, reduce demand by telling users to come back again later. This had been ignored until this time, then was scheduled to be written. Wizard, in their inimitable fashion, didn't start work on the code for months. We were delivered the first attempt 3-4 weeks before "Busy Day" [testing? testing? how could we load-test that?]. And within a matter of days the "Busy Tone" triggered for the first time. They'd almost scuppered us. Wizard had ignored a very important part of the specification, the "Busy Tone" page needed to be low-impact: to be a static page, not CGI, and not to load the database. What we were delivered did both things... When raised with the developer, the response was "I wasn't told that". Apparently reading detailed specifications wasn't part of this job description or maybe it was never passed on.
  • The Executive were shown the real-time monitoring system and guided roughly through what the various metrics meant.
    • This, and the load-balancer, were the single most important technical improvements for Operations.
    • The Executive and senior managers were able to call up the system load charts any time they liked. On our "busy day", 31-May, they intensively monitored the system all day [it was low overhead and run on a non-production server]. They were able to reassure themselves "all is well" in real-time and without needing to interrupt any technical staff. This empowered and resassured the management team and allowed them to be better informed whilst spending a lot less time and effort in doing so. Everyone was very happy.
After starting, an application was assigned a unique number and a user password stored, probably not 'in clear', but stored as a hash, like the Unix passwd file.
As well, a random number was generated that was stored as a hidden field in the HTML form.
This security device prevented URL-hacking, as happened to confidential information on some other Government websites at the time.

 Oracle supports a virtual table called a 'Sequence' that is guaranteed to return a unique, monotonically increasing, sequence of numbers: handling race-conditions and real-time demands.

In March or April, there was a serious Application fault: a client had restarted their Application, but incorrectly entered their application number (transposed digits). Somehow the wrong password was accepted and they were able to look around the confidential information of another person/entity - a massive security failure.

This event showed 3 problems in the Wizard process:
  • They used a simple Change Request system suitable for Development, but not for Operations.
    • We were unable to classify the fault as "Severity 1" and insist on an immediate work-around, or call an Emergency Operations meeting and consider taking the site off-line.
    • Despite multiple attempts to escalate the fault, nothing happened.
    • In the longer term we requested changes directed to addressing Operational issues properly, even getting access to their Fault Reporting system, but nothing came of it.
  • The fault had occurred before a public holiday and a weekend. Wizard did nothing to resolve this most serious fault for most of a week.
  • Over the weekend, without knowing anything of the Application or source-code, I guessed that the fault might be related to a race-condition in the initial page. Someone 'reconnecting' to an already open session would always get 'password OK'. It took 5-10 minutes to confirm this thesis on Monday morning and inform Wizard.
    • When given a reproducible test, Wizard were able to fix their code and deliver an update that day.
    • The Wizard team leader volunteered during the next weekly meeting, "You only beat us finding the fault by 20 minutes". That statement defies logical explanation... How could they have know when their testing or troubleshooting process would've led them to this insight?
The Wizard scripting language was an on-going source of concern.

One of the later problems occurred when we asked for a copy of the current source code. The updated versions that Wizard loaded onto the Systems were a "compiled" version and we were not given a compiler or any of the build tools.

Following around a month of disagreement by the (one of many!) Wizard Project Manager, including an attempt to first deny us access to 'their' Intellectual Property and then an attempt to charge us for what we'd already paid for, someone managed to get Wizard to read the (Govt. standard) contract which assigned Copyright and all Intellectual Property to the Client, BEP. Wizard had been duty bound to supply us the source code, but had not. When asked to rectify the problem, they'd been combative and recalcitrant.

Wizard had disbanded the ABN development team and were not interested in addressing performance concerns. Following a lot of cajoling, we got around a 20-fold performance improvement out of the scripting engine and simple caching was added to the Application.

The 'performance turning' was instead of doing an initial SQL select then reading all records (400-500 per application) ONE by ONe with a 'cursor'  - each read was a read/write transaction to the database, through the internal firewall.
That network and SQL processing load alone would've reduced our processing capacity 20-50 fold, but I never did test that.

After we'd pounded on them for a few months about this and other problems, Wizard very proudly came in one day and announced they'd implemented this massive speed-up in the Application. They'd issued a single SQL "select *" statement when the Application started and loaded the full set of records returned into a buffer for later use in a single transaction. This reduced the network/DB load to a single read/write, which made it possible for us to serve the load we finally handled on our "Busy Day".

In 1999, we had 300Mhz SPARC processors that weren't optimised to start large numbers of new processes - but were very good at handling multiple threads within processes.

Wizard were unable or unwilling to re-architect their code to suit the  SPARC and use threads.

The WhizKids made some stunning statements underlining their poor grasp of good Software Engineering and complete lack of Operational experience:
  • Forget about handling errors. [In response to a shell script I was crafting]
  • We won't do that, we store everything in the database. [Project manager commenting a code that was written and moving into production.]
  • Why isn't our test site the same as Production? [Really, do you have to ask? It does a whole bunch more things and is a bunch more expensive and larger.]
  • Why would I waste my time doing that for you? [yep, a Project Manager refusing to organise a log book]
  • We did that for performance reasons... [Writing two copies of various DB log files into different places on the same RAID filesystem.]
  • We've tested the code and it's OK. How does recompiling it with new parameters change any of that? [They'd hard-coded path names and Database constants into the code, rather than abstract them into a control or configuration file. They couldn't be convinced that once a binary was tested and accepted, then that exact set of bits was the 'configuration item', not a recompiled version.] 
Although Wizard proudly displayed their ISO 9001 accreditation and everything they produced had those Change Control and Release pages, they had almost no understanding of good quality processes.

More than once they ambushed us with "the bi-monthly major update depends entirely on this new piece of unapproved code being added within 5 days". This did not go down well in a high-security, highly controlled environment where every change had to be carefully inspected, justified and throughly tested.

The other "trick" they pulled was: "you signed off/accepted that. It's not our problem."
Using "Quality Systems" as a "Not My Problem"... [And pay us more money, and more, and more to attempt to get it right.]

Every two months, IIRC, Wizard did a major update of their systems.
They were supposed to be fully tested on their in-house test systems we'd supplied.
What they supplied was a bunch of tapes, very long sets of (paper) instructions and commands and a team of people. These 'upgrades' started with all systems going offline around 9AM and coming back 4-8 hours later. Every upgrade needed a frantic dash back to their office and extra materials/files brought across, they never went perfectly or quickly.

The commands were typed on the console of the machines as "super-user" ('god' or 'root').
I eventually refused to be part of these upgrades as they were insecure, unreliable and dangerous.
The WhizKids co-opted a non-techncal IT co-ordinator as the "BEP representative" responsible for overseeing the operation and for authorising the new site "go live". They were logged in as "super-user" to all production systems by the SGE staff.

That day I got a call at home, "how do we recover system libraries we've deleted?"

They'd blindly typed "rm -rf ./" as "super-user" into a command window without confirming the system they were on, or directory they were in.
This is pretty much the right command to destroy a whole system - which they almost did...

Early on I asked for a detailed set of actions for each role/person for a major upgrade so that it could be co-ordinated. They failed to disclose they already had extensive documentation and instructions and charged for the same document, slightly reordered.

I got an incredulous look when I asked Wizard to install the changes into a PreProduction environment that could be tested beforehand, and to leave the Previous Working Version available as a failback. With the load-balancer, we could've implement changes at 09:00 on a Monday, if a proper multi-version release system had been provided.

For high-security production systems, this should be a minimum.
Wizard, although they had extensive test systems, were unable to package their changes in a simple, reversible and auditable manner.

Competent admins would've both been able to derive exactly the files changes between releases, and have packaged the changes as a set of files and, if needed, scripts to install them and configure the systems if needed.

The last time I'd seen this level of incompetent chaos where every system upgrade required a further upgrade or changes of configuration items, was at Customs with "Management Solutions" of Hall and their FINEST accounting/financial management system. They went bust before Wizard.
The big omission in the Wizard solution was Performance: instrumentation, monitoring, reporting and Capacity planning were all missing.

As mentioned above, I designed and implemented a lightweight performance graphing system that proved very useful.

When I first arrived at BEP (September 1999?), I suspected that the design might be prone to failure under excessive load and did some load forecasting based on a couple of months data, but using some modelling and analysis techniques I'd recently learned.

The results concerned my boss and I: my predictions were for ~22,000 registrations on the final day.
That we were within 5% of this figure was sheer luck. It could've easily been 30% higher or lower.

When BEP approached Wizard for a paper on the performance of their system, this was produced:
  • "our target throughput is 1200 transactions a day, spread evenly over 24 hours. Or 50 transactions an hour. Our test system performs at that level".
  • they ignored extra load from abandoned or restarted applications, or users editing pages.
  • they ignored the additional load from system processes to transmit and track applications to the ATO gateway
  • they ignored the multiple firewalls between the web and database servers
  • And they failed to account for the difference between the test and production databases. [by a factor of 100-1000]
Wizard were unconcerned and unresponsive when we came up with a projection of 20 times their design maximum and that their system design guaranteed a catastrophic collapse when this occurred, causing severe embarrassment for the ATO, BEP and the Treasurer and Prime Minister.

My boss spent considerable effort in addressing aspects of this problem:
  • he hired Oracle to perform some real database performance tuning
  • he hired a firm in Sydney to analyse our daily load figures and model our response to load. This would tell us when the system was reaching capacity.
  • he approached SUN about extra capacity, disks and even machines
  • he got Wizard to implement multiple speeds-ups of their code.
Oracle identified one simple and important change for 'tuning'. The database server was spending half its CPU time parsing SQL statements - if Wizard chose to preparse their SQL code, it would've doubled the capacity of our backend. But they wouldn't change their code.

As mentioned above, we caught a break when the system overloaded at the start of March. It was the catalyst that allowed us to proceed with Busy Tone, purchase a substantial CPU upgrade and pressure Wizard to continue improving their software.

The absolute best piece of idiocy in the Wizard design/implementation was the database design.

Since Codd and Date came up with the theory behind Relational Databases and SQL, standard first steps with any database design is reducing it to "third normal form": this is the most compact and computational efficient database schema.

What did the WhizKids leave us?? Nothing like a good DB design.

It consisted of just 3 fields, the primary key (a number and a text 'key') and a value stored in a 255 long "varchar". This might be good for a quick hack or small test, but for a $1MM project, it's not just ridiculous, but as close to professional negligence as I've seen.

The 'key' was some sort of identifier for the various field names on the forms, plus all the system states.
Heaven knows how they dreamt up and kept a master list of all those names. My estimate was 400-500 fields. Perhaps they concatenated standard field names with the originating form, even creating two sub-fields.

A reasonable estimate for the space needed for all data for one Application might be 8-12Kb.
Instead, we got 500 lots of 256 bytes, or ~75Kb.

The killer though was the number of records in the single monster database table.
Instead of 2-8 short records per Application, we had 400-500.

The final database for 600,000 applications had over 250 Million records, consuming up to 30Gb (massive at the time), instead of 1-2M records and 3-5Gb.

But it gets worse...

There were two numbers generated by the Application for every new client application: a 'sequence' (an integer) and a random number.

Which should be used in the primary key? The sequence, that's part of what its designed for.
Databases can optimise indexes for sequential keys, especially integers, very well. For hybrid keys, integers and text fields, its not nearly as efficient.

Wizard used the random as the primary key of the database, hybridised with a text string.
This is such bad practice it defies belief.
It maximises the work that the database has to do to create an index and guarantees worst-case lookup performance.

Why would any knowledgable, competent professional do that? I can't guess.


But it gets worse...

The ATO needed client applications electronically transferred to a front-end system, where they were copied to tape and transferred to the secure back-end processing system. When the records had been loaded, an acknowledgement was sent back to the ABN system to mark the client application as complete and lodged.

The client application was written into a standard interchange format (I never looked into what they used. Could've been XML, could've been structured text, could've been an encrypted binary format).

These messages were sent via email, package in a standard "message digest" format with an included digital signature (MD5 hash?) for end-end checking.

As a backup, in case the email system failed, the messages could be dumped to files, that were then written to CD-ROM and imported directed to the ATO back-end system.

The files were named for the sequence number, the email used the random as the message number.

When we created a test disk and tried to cross-check and verify, we found it impossible.

But it gets worse...

The ATO had carefully specified a message digest format for the emails.
This allows multiple files to be included in a single email message.
Wizard only included one file in each email message.

At the time, 20,000 messages being sent during the day was a reasonable load on the SGE.
The SGE handled the email traffic for a large number of Agencies.

At one point, the process, it ran on the database server, that selected emails and sent them to the ATO front-end, failed and this went undetected for a day or two without raising an alarm or being noticed.

Someone from Wizard visited the datacentre and noticed, they "released the queue" (another bone of contention. They failed to inform us of visits, the work they performed when there or to document it sufficiently so it could be duplicated).

At which point, in the middle of the day, it flooded the SGE mail system with 5-10,000 messages, creating a 2-3 hour delay for all email for all Agencies using the SGE.

Wizard implemented some 'fix'.
Then promptly repeated the error... We were NOT popular with the SGE folks.

There must have been more, but that's quite enough.

And what happened next?

  • We survived "Rush Hour". The system hit 100% before 9AM and stayed maxxed-out until I went home at midnight.
    • The back-end filesystem was overloaded, something I hadn't anticipated, nor instrumented.
    • Part of the effect of this was DNS requests were being lost, causing email to fail.
  • David and Charles built a replacement system for $100,000 total in 3 months.
    • duplicated hardware, firewalls, load balancer, database, licenses, contractor costs.
    • Like everything with Wizard, 10-times smaller, cheaper and probably faster.
  • I tried to interest the ATO project manager in writing up what we'd learnt as the first major Government on-line transaction.
    • My contract was terminated before it got anywhere and years later I followed up and nothing had ever been published.
  • In 2006 I found a fragment of the process accounting data and reanalysed it, resulting in the CMG-A paper.
    • In that follow up study, I understood how close we'd come to a secondary failure: saturation of the database by the Busy Tone pages.