Saturday, March 5, 2011

NetApp Deduplication volume size limitations

I found this in one of the NetApp documentations for maximum deduplication volume size limits.

Create a new aggregate while zeroing spare disks

cherry-top# aggr create aggr1 -r 14 -d 0a.16 0a.17 0a.19 0a.22 0a.23 0a.24 0a.25 0a.26 0a.27 0a.28 0a.29 0a.32 0a.33 0a.34
aggregate has been created with 11 disks added to the aggregate. 3 more disks need
to be zeroed before addition to the aggregate. The process has been initiated
and you will be notified via the system log as the remaining disks are added.
Note however, that if system reboots before the disk zeroing is complete, the
volume won't exist.


cherry-top#  vol status -s


Spare disks
RAID Disk Device HA SHELF BAY CHAN Pool Type RPM Used (MB/blks) Phys (MB/blks)
--------- ------ ------------- ---- ---- ---- ----- -------------- --------------
Spare disks for block or zoned checksum traditional volumes or aggregates
spare 0a.35 0a 2 3 FC:A - ATA 7200 635555/1301618176 635858/1302238304 (zeroing, 2% done)
spare 0a.36 0a 2 4 FC:A - ATA 7200 635555/1301618176 635858/1302238304 (zeroing, 2% done)
spare 0a.37 0a 2 5 FC:A - ATA 7200 635555/1301618176 635858/1302238304 (zeroing, 2% done)
spare 0a.38 0a 2 6 FC:A - ATA 7200 635555/1301618176 635858/1302238304 (zeroing, 2% done)
spare 0a.39 0a 2 7 FC:A - ATA 7200 635555/1301618176 635858/1302238304 (zeroing, 2% done)
spare 0a.40 0a 2 8 FC:A - ATA 7200 635555/1301618176 635858/1302238304 (zeroing, 2% done)
spare 0a.41 0a 2 9 FC:A - ATA 7200 635555/1301618176 635858/1302238304 (zeroing, 2% done)
spare 0a.42 0a 2 10 FC:A - ATA 7200 635555/1301618176 635858/1302238304 (zeroing, 2% done)
spare 0a.43 0a 2 11 FC:A - ATA 7200 635555/1301618176 635858/1302238304 (zeroing, 2% done)
spare 0a.44 0a 2 12 FC:A - ATA 7200 635555/1301618176 635858/1302238304 (zeroing, 2% done)
spare 0a.45 0a 2 13 FC:A - ATA 7200 635555/1301618176 635858/1302238304 (zeroing, 2% done)


cherry-top# aggr status -v
Aggr State Status Options
aggr1 creating raid_dp, aggr nosnap=off, raidtype=raid_dp,
initializing raidsize=14,
ignore_inconsistent=off,
snapmirrored=off,
resyncsnaptime=60,
fs_size_fixed=off,
snapshot_autodelete=off,
lost_write_protect=off
Volumes:

Plex /aggr1/plex0: offline, empty, active




Friday, March 4, 2011

Cluster giveback cancelled or waiting

Cluster node is down and waiting for giveback. Trying to perform a cf giveback from the partner which has taken over generates the following error message:

filer(takeover)> cf giveback
filer(takeover)> Thu Dec 21 21:01:55 EST [filer (takeover): cf_main:error]: Backup/restore services: There are active backup/restore sessions on the partner.
Thu Dec 21 21:01:55 EST [filer (takeover): cf.misc.operatorGiveback:info]:
Cluster monitor: giveback initiated by operator
Thu Dec 21 21:01:55 EST [filer (takeover): snapmirror.givebackCancel:error]: SnapMirror currently transferring or in-sync, cancelling giveback.
Thu Dec 21 21:01:55 EST [filer (takeover): cf.rsrc.givebackVeto:error]: Cluster monitor: snapmirror: giveback cancelled due to active state
Thu Dec 21 21:01:55 EST [filer (takeover): cf.rsrc.givebackVeto:error]: Cluster monitor: dump/restore: giveback cancelled due to active state
Thu Dec 21 21:01:55 EST [filer (takeover): cf.fm.givebackCancelled:warning]: Cluster monitor: giveback cancelled


To resolve, make sure the  following outputs has no existing relationships or active sessions.

# snapmirror status -l
# snapvault status 
# ndmpd status (kill any active sessions)

Issue cf giveback on the system after confirming all the above steps are verified.

Sunday, February 20, 2011

World’s first shipping 7200 RPM 3TB enterprise-class HDD from Hitachi

The Hitachi Ultrastar™ 7K3000 is the world’s first and only 7200 RPM hard drive rated at 2.0 million hours MTBF and backed by a five-year limited warranty. The Ultrastar 7K3000 represents the fifth-generation Hitachi 5-platter mechanical design, first introduced in 2004, and has been field proven by top server and storage OEMs as well as leading Internet giants. When the highest quality and reliability are a top requirement, customer field data proves that the Ultrastar 7K3000 delivers by reducing downtime, eliminating service calls and keeping TCO to a minimum. Engineered for the highest reliability, the Ultrastar 7K3000 is not only put through grueling design tests during development but must also pass stringent ongoing reliability testing during manufacturing. Across the entire Ultrastar family, world-class quality control, combined with scientific root-cause analysis and multi-faceted corrective actions, ensure that Hitachi GST remains the recognized leader in quality and reliability for enterprise-class hard drives.

Highlights:

• 2.0 million hours MTBF
• Up to 3 terabytes of capacity
• 6Gb/s SATA and 6Gb/s SAS models for configuration flexibility
• Dual Stage Actuator (DSA) and Enhanced Rotational Vibration Safeguard (RVS) for robust performance in multi-drive environments
• 24x7 accessibility for enterprise-class, capacity-optimized applications
• 5-year limited warranty

Product Documentation : http://www.hitachigst.com/tech/techlib.nsf/products/Ultrastar_7K3000

Saturday, February 12, 2011

Minimum size for root FlexVol volumes


Storage system model   Minimum root FlexVol volume size

FAS250                                        9 GB

FAS270                                       10 GB

FAS920                                       12 GB

FAS940                                       14 GB

FAS960                                      19 GB

FAS980                                       23 GB

FAS3020                                    12 GB

FAS3050                                    16 GB

F87                                              8 GB

F810                                            9 GB

F825                                           10 GB

F840                                            13 GB

F880                                            13 GB

R100-12TB                                 13 GB

R100-24TB                                 19 GB

R100-48TB                                 30 GB

R100-96TB                                 53 GB

R150                                           19 GB

R200                                           19 GB

Thursday, January 27, 2011

All about /etc/rc file

The /etc/rc file contains commands that the storage system executes at boot time to configure the system.

What startup commands do  

Startup commands are placed into the /etc/rc file automatically after you run the setup command or the Setup Wizard.

Commands in the /etc/rc file configure the storage system to
1) Communicate on your network 
2) Use the NIS and DNS services 
3) Save the core dump that might exist if the storage system panicked before it was booted

Default /etc/rc file contents

To understand the commands used in the /etc/rc file on the root volume, examine the following sample /etc/rc file, which contains default startup commands:

#Auto-generated /etc/rc 


hostname filerA
ifconfig e0 `hostname`-0
ifconfig e1 `hostname`-1
ifconfig a0 `hostname`-a0
ifconfig a1 `hostname`-a1
route add default MyRouterBox
routed on
savecore


Explanation of default /etc/rc contents 

Description : hostname filerA
Sets the storage system host name to "filerA."

Description : 
ifconfig e0 `hostname`-0 
ifconfig e1 `hostname`-1
ifconfig a0 `hostname`-a0
ifconfig a1 `hostname`-a1
Sets the IP addresses for the storage system network interfaces with a default network mask.
The arguments in single backquotes expand to "filerA" if you specify "filerA" as the host name during setup. The actual IP addresses are obtained from the /etc/hosts file on the storage system root volume. If you prefer to have the actual IP addresses in the /etc/rc file, you can enter IP addresses directly in /etc/rc on the root volume.

Description: route add default MyRouterBox

Specifies the default router. You can set static routes for the storage system by adding route commands to the /etc/rc file. The network address for MyRouterBox must be in /etc/hosts on the root volume.
 
Description : routed on

Starts the routing daemon.
 
 
Description: savecore
Saves the core file from a system panic, if any, in the /etc/crash directory on the root volume. Core files are created only during the first boot after a system panic.

Friday, January 14, 2011

BaseBoard Management Controller (BMC) setup

You can manage your storage system locally from an Ethernet connection by using any network interface. However, to manage your storage system remotely, the system should have a Remote LAN Module (RLM) or Baseboard Management Controller (BMC). These provide remote platform management capabilities, including remote access, monitoring, troubleshooting, and alerting features.

cherrytop# bmc setup

The Baseboard Managment Controller (BMC) provides remote managment capabilities
including console redirection, logging and power control.
It also extends autosupport by sending down filer event alerts

Would you like to configure the BMC? (y/n)? y
Would you like to enable DHCP on BMC LAN interface? (y/n)? n
Please enter the IP address for the BMC [0.0.0.0]: x.x.x.x
Please enter the netmask for the BMC [0.0.0.0]: x.x.x.x
Please enter the IP address for the BMC gateway [0.0.0.0]: x.x.x.x
Please enter the gratuitous ARP Interval for the BMC [10 sec (max 60)]:

The BMC is setup successfully.

The following commands are available; for more information
type "bmc help "
bmc help 
bmc setup 
bmc status 
bmc test
bmc reboot

This can be done online and is transparent to the Filer/servers connected to the NetApp Array. 

 

Monday, December 20, 2010

How to expand an A- SIS enabled volume that is nearing the "vol size" limit

WARNING: If this is being performed to free up space to bring a LUN back online, another method of cleaning up space should be considered, such as deleting snapshots and disabling automatic snapshots until a maintenance window can be scheduled as the 'undo' process can take a significant amount of time.

To increase the size of an A-SIS enabled volume beyond the maximum limit for A-SIS, the A-SIS service must be turned off and the changes undone. Undoing A-SIS will re-inflate the file system and could require more disk space than is available in the A-SIS enabled volume. There is no way to expand the volume size until the undo is completed, so the recommended course of action is to create and use a temporary volume and migrate data necessary to free enough space for the re-inflation to complete.
WARNING: Once the volume is grown beyond the maximum size supported for A-SIS, A-SIS will be disabled.

WARNING: Disabling A-SIS will require additional disk space as files will be undeduplicated.

WARNING: Using "sis undo" may require rebaselining of snapmirror or snapvault relationships.

Complete the following steps to undo A-SIS:


Note: The undo must be performed from diag mode. The sis undo command can take some time (hours) based on how much data is being un-deduped and the filer type.

1- Enter df -s
Note the space saved, this is the amount of space that will be necessary for the re-inflation.
2 - Enter df
Note the available space. If it is not greater than or equal to the space saved found in the previous output, space will need to be cleared using other methods to complete the undo (e.g., deleting Snapshots or migrating data).
3 - Enter sis off
4 - Enter priv set diag
5 - Enter sis undo


Once the undo is complete, the volume will be a normal FlexVol volume that can be expanded.

When trying to access the filer using a NetBIOS alias, error message: Decrypt integrity check failed

When trying to access the filer's NetBIOS alias, the following error messages are generated:


[auth.trace.authenticateUser.krbReject:info]: AUTH: Login attempt from 10.20.1.13 rejected by Kerberos.
[cifs.trace.GSSinfo:info]: AUTH: notice- Could not authenticate user.
[cifs.trace.GSSinfo:info]: AUTH: notice- Decrypt integrity check failed.


Issue
The Active Directory had a stale computer account that had the same name as the NetBIOS alias used to contact the filer. The NetBIOS alias may have been created in the Active Directory during vFiler testing, but the account had not been removed.


Check the following:
1. Check for another account in the same AD forest that has the same name as the filer. This can either be a stale account (left over from a previous situation), or another machine.


2. The reason you can connect when you specify an IP address is that the client uses NTLM instead of Kerberos in that situation. When the client gets a Kerberos ticket, it is most likely getting a ticket for the wrong machine.


Friday, December 3, 2010

How to Reboot / Reset an HP Blade iLO2

Every once and awhile an iLO2 Remote Control session will get "stuck". By this I mean when you try to connect, it will say that there is a session already in progress. A reboot of the iLO2 will clear this state. To restart the iLO2, follow these steps.

* From the Web Administration page for the iLO2, you should be on the System Status tab by default 
* Click Diagnostics on the left hand side
* On the bottom of this page is a Reset button for the iLO2