Sunday, 8 March 2009

Still here, still managing Backup Exec every day...

It's been some months since we've postted here, and the foolish (and those who never actually deal with Backup Exec) would probably believe that's because we've eithe forgotten about the blog, or we've got everything working.

You'd be wrong.

It's true to say that we've got a little more success, and now have 5 10d boxes and 1 12d box, all running, and, most of the time it tends to play well. Which is lovely, but when it does go wrong, it tends to lose it completely.

Here are a few problems we've currently got:

a) An old "Managed Media Server" that is long since departed just won't go away from all parts of the UI - most of it knows it has gone, but some parts still show it, as if it may somehow come back one day. It won't.

b) 2-3 Jobs are stuck in an external status where they're on "On Hold, Running" according to the status. That's not true. In fact, they've been stuck there for a year. Meaning I can't delete the now-redundant Policies, Templates or Selection Lists for those jobs. They're just stuck there forever.

c) Sometimes a job fails claiming the cause to be a Communications Failure. Communications Failure is Backup Exec speak for "most problems". There is naff all wrong with the communications, and normally we resolve it by deleting/recreating the job.

d) Synthetic jobs, well let's see. They suck. They only work in exacting circumstances, and the minute you step out of line of one of those or a job is missed, well that's your life made hell. They start failing, come up with lots of silly errors and you end up re-creating them, waiting for a full again. So you tend to not bother, and just do an old-style Full/Incremental set, since they normally work.

In the case of 12d, it does tend to be a little better, particularly with Exchange Backups. Except it STILL doesn't properly manage media, so the IMG foldes it creates don't always get deleted (although it reckons they will). Still no joy on having the B2D Files reused or deleted. Hell no, that'd make sense.

So yeah, I'm still here, managing Backup Exec, 7 days a week, doing what it should do for it, and going mental every time I come in to find it's just collapsed without warning. Quality it is not.

Sadly I've still yet to find a better solution at a price point that is sane.

Friday, 31 October 2008

Why are error messages not unique

That's what I want to know.

Why is it that you get an error message in Backup Exec, and it spits out an error for you, and so you click on it, which takes you to a web page with a description of that error, right?

Wrong. With Symantec Backup Exec, it just takes you to a list of issues which may or may not be remotely close to what you have issues with, and rarely has any useful answer.

Here I am today again trying to find out why certain jobs keep failing without any sort of sane error reason. Another day fighting Backup Exec.

Wednesday, 17 September 2008

Server Paused. In a not-actually-paused-at-all kinda way

For the past 2-3 weeks our CASO Backup Exec Server has been INSISTING that the "Job Status" for every single job is "Server Paused". While it isn't actually uncommon for that to be the case, and certainly I've seen a server have this status from time to time, it isn't the case for days and not on all of our media servers.

It seems the answer is simple (but annoying)... here's what a Symantec Forum post says:

----------------
I have had this problem many times. Like a plague. The fix most times is simple. Go to the devices tab. Highlght your media server. Right click and pause it. Then unpause it. This is caused by an interruption which corrupts a file. Backup Exec shows nothing paused. Just fixed this today after 10 days following a problem with a tape drive. Symantec support knows about it but hasn't published the fix.
----------------

Friday, 12 September 2008

A few months later...

Backup Exec 10d continues to offer variable results. For straightforward jobs it tends to do a reasonable job, most of the time.

However, each time we add a set of Synthetic jobs to the mix, things start going horribly wrong, and the backup exec engine regularly dies.

While I currently refuse to pay any money for upgrades since they promised these features were part of 10d, and they simply don't work, I have had more luck with an eval of 12d. Of course until we've been running it as long as 10d and it has the same sort of load we won't know for sure, but I don't hold out much hope!

Tuesday, 26 February 2008

Event numbers

So, now that we've managed *touch wood* to iron out the big issues that have been plaguing us for ages, perhaps it's time to look at some of the more annoying niggles we have to contend with.

I've always thought that many event numbers produced by most applications were simply too short and simple, after all, how can you hope to cover all the eventualities with only 4/5 characters. With the change in Backup Exec in recent versions to much longer event numbers you would think they could tie things down much more specifically, but apparently not.

Looking at an error we've been seeing recently, we're getting this when running a backup :

V-79-57344-33938 - An error occurred on a query to database .

So you might think with a possible 1,000,000,000,000 event numbers available that this error would be specific to the problem I'm seeing, and allow me to find more information directly relating to it... of course you'd be wrong.

Clicking on the link provided by the Job History shows me 9 different articles all apparently relating to this one error. 6 of them refer to restore jobs rather than backup jobs, 2 relate to SQL 2005 not 2000 as is the case in this instance, 2 relate to running SQL 7 alongside 2000 which we're not doing and of the 2 which do relate to backups in their titles, the contents either refer to SQL 7 or to you having an incomplete restore operation preventing the backup from running.

Now I realise that producing written content for up to a trillion error numbers would be a mammoth task, but personally if I'm searching for a specific event number and no content exists for it yet I'd prefer to know that, rather than trawl through mountains of knowledge base articles in the hope that one might actually be relevant to my situation! I can always then search using words from the error message to find articles that are similar, and try things from there!

Monday, 14 January 2008

Reliability? Can we say we're onto Stage 2?

I'm rather scared, and pleased, to say that we appear to have genuinely made it work - in my last few posts I was still sceptical that Backup Exec was just being nice to us, but it does appear it is genuinely now acting like a competent backup product.

One server has now been up for 27 days, and is still running backups quite happily, with 2,000 odd jobs run in that time (66 a day or so), and the CASO server just getting on with it's jobs and delegation. No more tears.

Our CPS system is still working - that's the most flawless part of Backup Exec I've seen. It's absolutely awesome - it says continuous protection, and it offers just that. We installed it, got over one hurdle of making it listen on a specific IP, and that was that. Our main file servers have just-below-real-time backups 24/7.

Our next challenge (Stage 2, which took 18 months to get to!) is to build a comprehensive reporting and restore testing infrastructure around the software. We want to know everything possible about what it does, so we can provide internal quality reports, ensure we meet SLAs and finally, but not least, ensure customers can be assured of regular, reliable backups.

Restore testing is often overlooked, but certainly not here! It is an essential, and core part of our plan to operate regular (hopefully scheduled) restore tests so we can be sure those backups actually work - as someone else we know found out to great cost - backing up isn't enough!

Saturday, 29 December 2007

Forgotten about this blog? Not at all

But the problem is Backup Exec does appear to largely work reliably now - at the moment we're on 12 days 5 hours. It's hardly something you can call "reliable to 5 9's (e.g. 99.999% reliable), but it's certainly better than every few hours!

Just to recap the top tips...

1) Don't set more than a handful of jobs to run at once. They won't.

2) Although policies and templates are good ideas, expect to create lots of them because of tip 1.

3) Have plenty of Disk to Disk Folders, and set no more than 1 or 2 jobs per set as the concurrent limit.

4) Make good use of the Disk Reserve settings to ensure it has no excuse to run out of space and fail your jobs for weeks.

5) Set sane retention policies (e.g. don't just set it to weeks for the sake of it)

6) Split jobs by location, server, and drive.

7) Backup Exchange or SQL (and other "Agent" options) as distinct jobs

8) Keep your eye on Backup Exec at all times.

Friday, 23 November 2007

And now, for CPS...

Something truely amazing has started happening, as Backup Exec is *still* running, 6 days in, and it looks like we may have turned a corner with it in terms of making it think, act and work, like a competent backup solution.

As a result we've been able to look at other things - since my colleague tried the "CPS" or Continuous Protection Server part of Backup Exec some time ago, we found it was pretty good stuff - it worked straight out of the box, so we figured we'd give it a try - after all it would be more than a little useful if we could just replicate data from one site to another as part of a backup system, as it would boost our protection against failure of one of our key servers.

Installation of the CPS Server was easy, and worked first time. However, getting it to talk to our remote server was a little more difficult, as it's on a different subnet, and, as a bonus, the CPS Server also happens to have multiple NICs on different networks. To annoy us, CPS decided to pick the wrong IP to bind to (and there's no indication it will do this, nor any GUI to choose it).

It's OK though, easily fixed - with a famous registry edit to set a "PreferredAddress" on the CPS Server, and, on the CPS Machine being copied, changing it's "Gateway" address for the Veritas Software to be the IP of the CPS machine and not it's host name (even though it resolves to the same).

Right now we have a working CPS job doing the initial copy of our data, after which it should continuously protect... rock on.

Tuesday, 20 November 2007

I'm still in a dream

I must be. Because we've still got 3 working Backup Exec Media Servers, CASO is working, backups are working, and, with a couple exceptions they all run.

We've got 2 troublesome jobs which fail as they just have issues spitting the files out to the Media Servers and we'll work on those, and one that just fails on System State every day despite Windows itself being able to back it up, Backup Exec insists that "a failure occurred reading an object" and later claims that "shadow?copy?components" is a corrupt file.

So right now, we've had 3 days and 14 hours of running backups - you know we may even get to run test restores because it's not crashed.

Sunday, 18 November 2007

Did we hit April 1st? What has been done with the Real Backup Exec?

In a brave move, given the near 3 day successful run with Backup Exec over last week, and given that I was more interested in seeing Bill Bailey live over the weekend than staring into the "Alerts" list on Backup Exec, I figured we'd just see if it could cope again.

Amazingly, it has. So far. It's now late Sunday evening but everything is running, and, to top it I've got 3 media servers online, running, and doing what they're designed for. We've even found time to expand the storage on the smallest of the servers from around 1Tb to 3Tb ready for when it starts to get a real workload.

Friday, 16 November 2007

It was too good to be true.

Backups have run for 2 days, 23 hours, and just crashed I actually thought we'd make 3 days then.

It appears that "BECatSrv.dll" decided to give up and die.

Now to see if it recovers...

Thursday, 15 November 2007

"d" is for "tape"

Something is wrong. Very wrong. Backup Exec is still working. That's over a day. It's not crashed or had a funny 10 minutes where it doesn't work. None of this "server paused" rubbish. Still meanwhile in the world of the admin's maintaining it, we've had a good old debate about how the whole thing is supposed to work, how symantec interpret "tape" backups when they're actually disk backups.

The only thing we did resolve is that the "d" in Backup Exec 10d means "tape". Yes. it means Tape. Which is probably how Symantec would have preferred we kept things.

Eye Tests

I'm starting to think I need an eye test. In our office I'm the only one with good eyesight that doesn't already wear glasses. Amazingly though, I think that may have to change.

My rationale is that this morning I've opened the old Backup Exec status window, yet there isn't a Server Paused, Loading Media, Queued or other failure at all, and over the past 24 hours we have just 13 out of 250 odd jobs failed, and the 13 are largely "missed" backups which just mean I need to refine the windows so there is enough time for them all to run (I never get that far normally as it crashes so everything's "missed".

So yeah, either my eyes are playing tricks on me, or Backup Exec has performed a minor miracle.

Wednesday, 14 November 2007

Making it work: Tip #2

One of my pet hates of Backup Exec is it's rather poor disk space management for Backup To Disk Folders. There are 2 approaches to Disk Space Management, the Sane Way, and the Stupid Way.

Guess which Backup Exec chooses...

As you can't set a limit on the amount of space a Backup To Disk volume can occupy (almost like a LTO-1 tape has a 100Gb storage capacity, and that's that, if you run out, you run out), you'd expect to be able to do similar on Backup Exec. But while you can tell it how big an individual Backup To Disk FILE is, you can't set a limit on the total space a "Disk" can occupy. The limit of course is the physical space on the drive itself.

The "workaround" is that if you had a drive, let's say it's 500Gb, and you have 6 Backup to Disk "Folders" on that drive. Set each one to have a Disk "Reserve" of something (the default is "nothing", naturally), let's say 10Gb. When you first use Backup Exec, it'll just keep making "B2Dnnnnn" files, but once you get to <10Gb left on the PHYSICAL drive, the Backup Exec Folders will then theoretically start to be overwritten, preventing the problem of running out of space completely.

Follow that? No. OK, try this:

PHYSICAL DRIVE (we'll call it Drive E:) has 500Gb Space, and it's only being used by Backup Exec.

You have 6 folders, each set with 10Gb "Disk Reserve":

Folder A, Folder B etc

Roll forward 2 weeks...

Folder A has 40Gb
Folder B has 2Gb
Folder C has 19Gb
Folder D has 125Gb
Folder E has 1Gb
Folder F has 28Gb
Folder G has 6Gb
...and the physical drive has 279Gb Free.

Roll forward 4 weeks....

Folder A has 100Gb
Folder B has 10Gb
Folder C has 90Gb
Folder D has 160Gb
Folder E has 2Gb
Folder F has 118Gb
Folder G has 10Gb
...and the physical drive has 10Gb Free.

At some point between weeks 3 and 4, the physical space hit 10Gb, so old "media" was overwritten. That's the theory of how this works.

Here's the reality:

* Make sure the PHYSICAL storage is MORE than the expected DATA need - that's full backups, storage for daily changes, incrementals, synthetics and so on, plus any other data you have on the same volume (perhaps Catalogs). It pays to have perhaps 50% more storage than you truely need, and more if you can.

* Make sure your storage calculations are enough for the retention. So if you've got a 4 week retention, do a weekly full, you need x 4 plus incremental storage. Realistically, maybe 6 times the storage.

* If your overall data storage decreases, don't expect more space to appear on the phyiscal drive. Because it treats "media" like a tape, it handles it the same way. If you had 30 tapes, and you only needed 20 of them now, the other 10 don't get erased from the earth, they just stay kicking about in your cupboard. Backup exec keeps your virtual "tapes" in it's cupboard. (we'll explain how to manage this better some other day).

* Regularly monitor your Backup to Disk Folders, particularly once you're getting to the space limits, because if your reserve is too big, and the retention period longer than you can cope with in physical space terms, you'll start getting failed jobs if there is simply no overwritable media left to be used.

* Having "overwriteable" media available is handy in some ways, as it means data beyond the retention period can still be available if the media hasn't yet been overwritten, almost building in a "last chance saloon" retention period, but it is also likely to consume all available space.

* If not all data is equally important, consider having different backup to disk folders, so that more critical data can be given longer actual retention times, but don't forget overwriteable media in one folder can't be claimed by another to add space, so good planning of physical space, folder space and necessary space is required.

If you follow this guide you can reduce the misery of the lacklustre Backup Volume Management. It won't allow you to have absolute limits for a media set (which you'd no doubt want), but it does stop you just running out of space.

Uptime - 1 day!

A minor miracle has occured - the Backup Exec services have decided to play nicely together, and as a result the CASO server has now been running for 1 day. It sounds like it's nothing, the sort of uptime that would have perhaps made a Windows 98 use proud, but in fact, with Backup Exec, a whole day without something crashing is certainly a miracle.

Right now we have 3 servers, all up, all running jobs, all receiving jobs delegated from the CASO box, and barely any failures (a couple of those are genuinely not the fault of Backup Exec). It will never last, surely?

d means "dumb"

I think that's what it means. Dumb as in "poorly thought-out", "stupidly implemented" and so on. At least that's what I believe. It certainly doesn't point to any amazing disk capabilities as they don't work.

What's the point of this 'd' nonsense?

So the 'd' in 10d, is supposed to represent 'disk', as in you can now backup to disk, rather than tape, but hang on... we've had disk-based backups since v8.6, so what's the beef?

Well aparantly the emphasis is more on the 'staging' side of things, ie you can backup disk-to-disk-to-tape. Brilliant I thought, that's just what we need. We can do the first backup to disk, which is nice and fast, then we can put that media set onto tape and store it offsite, presumably using the new 'duplicate' feature. That way, we've got the data onsite in case we need a quick restore and a duplicate copy offsite for disaster recover - what more could we want? Even better, it'll all be really easy to do now, because we have 'Policies', 'Selection Lists' and 'Templates' to save time.

NO, NO, NO! What was I thinking?! Did I actually expect the new features to work, as PROMISED? Silly me. That was HOPING for too much. Skip forward to REALITY, which actually means 'disks' are treated exactly the same as tapes. Why? How much more effort would it have been to expect that a 'd' product would have at least some slight comprehension of disk-based media? I'm sorry, but I'd expect a little more intelligence, such as being able to quote the maximum disk size, so that I could actually allocate a quota for a particular device I was backing up. Too much to ask obviously.

And another thing... Am I the only person that thought that this new 'duplicate' media set feature would work between different Managed Media Servers? After all, they're connected via a LAN and they have the ability to share Files and Folders via normal, everyday UNC shares. You can even create a 'duplicate' job and set the source and target devices from different Media Servers - they're all there in the list, but when you try to submit it, it says no. Now that's something that would've been bloody useful. Imagine, you have a Managed Media Server on another site, via a VPN link and being able to automatically duplicate your Media Set to it, so you can relax in the comfort of knowing that you always have an offsite backup? How much more effort would that have been?

Obviously Symantec are very good at painting pictures, but not so good when it comes to implementation. Maybe they should get out of this software game and get into bull**** marketing instead.

A new morning, a new failure...

It's just another average morning, a little nippy outside, terrible morning TV still exists (when will it die?) and I've got a raft of Backup Exec failures on my hands.

It seems that one of our sites has a condition I'll call "fussity", whereby the CASO server will submit it the jobs it's supposed to do, and, for about 90% of them, it'll just spit the dummy. Throwing "loading media" at you for a random period (could be 2 mins, could be 20 mins) eventually it gives up and fails the job. Only for the next one to work. And then next 9 to fail, and so on.... the last time we had a case of "fussity" the only way to stop it being so picky about jobs was to completely uninstall, rename the server, reinstall, reservice pack, and then re-create your devices, media, jobs and so on. I'm hoping to avoid it this time.

Meanwhile the main site figured it wasn't in the mood initially last night until around 8:30pm when I rebooted it. It just kept "recovering" it's own jobs. Good stuff.

All in a day's admin of Backup Exec.

Tuesday, 13 November 2007

Backup Exec Upgrade Cycle Explained:




Here we attempt to explain the Symantec Backup Exec upgrade life cycle. Start at 'Promise' and work your way round clockwise.

A New Strategy

I'm going to try a new strategy. We're going to create new jobs, one by one, for the servers. Each job will backup using old style Full/Incremental jobs, to the local server with basic 4 week rotation. Not even synthetics now, we can kiss goodbye to even more space. Thanks Backup Exec!

Let's see in a few days if that works either (see the lifecycle post below) - I'm at the "hope" stage again here...)