Showing posts with label Database Server (SQL). Show all posts
Showing posts with label Database Server (SQL). Show all posts

Apr 15, 2009

File Systems in Detail

File Systems in Detail

What is a File System?

A file system is a method for storing and organizing computer files and the data they contain to make it easy to find and access them. File systems may use a data storage device such as a hard disk or CD-ROM and involve maintaining the physical location of the files, they might provide access to data on a file server by acting as clients for a network protocol, or they may be virtual and exist only as an access method for virtual data. It is distinguished from a directory service and registry. More formally, a file system is a special-purpose database for the storage, organization, manipulation, and retrieval of data.

Most file systems make use of an underlying data storage device that offers access to an array of fixed-size blocks, sometimes called sectors, generally a power of 2 in size. The file system software is responsible for organizing these sectors into files and directories, and keeping track of which sectors belong to which file and which are not being used. Most file systems address data in fixed-sized units called "clusters" or "blocks" which contain a certain number of disk sectors. This is the smallest amount of disk space that can be allocated to hold a file.

However, file systems need not make use of a storage device at all. A file system can be used to organize and represent access to any data, whether it be stored or dynamically generated.

File Names Conventions

Whether the file system has an underlying storage device or not, file systems typically have directories which associate file names with files, usually by connecting the file name to an index in a file allocation table of some sort, such as the FAT in a DOS file system, or an inode in a Unix-like file system. Directory structures may be flat, or allow hierarchies where directories may contain subdirectories. In some file systems, file names are structured, with special syntax for filename extensions and version numbers. In others, file names are simple strings, and per-file metadata is stored elsewhere.

Metadata

Other bookkeeping information is typically associated with each file within a file system. The length of the data contained in a file may be stored as the number of blocks allocated for the file or as an exact byte count. The time that the file was last modified may be stored as the file's timestamp. Some file systems also store the file creation time, the time it was last accessed, and the time that the file's meta-data was changed. Other information can include the file's device type, its owner user-ID and group-ID, and its access permission settings.

Arbitrary attributes can be associated on advanced file systems, such as XFS, ext2/ext3, some versions of UFS, and HFS+, using extended file attributes. This feature is implemented in the kernels of Linux, FreeBSD and Mac OS X operating systems, and allows metadata to be associated with the file at the file system level. This, for example, could be the author of a document, the character encoding of a plain-text document, or a checksum.

Hierarchical file systems

The hierarchical file system was an early research interest of Dennis Ritchie of Unix fame; previous implementations were restricted to only a few levels, notably the IBM implementations, even of their early databases like IMS. After the success of Unix, Ritchie extended the file system concept to every object in his later operating system developments, such as Plan 9 and Inferno.

Facilities

Traditional file systems offer facilities to create, move and delete both files and directories. They lack facilities to create additional links to a directory (hard links in Unix), rename parent links, and create bidirectional links to files.

Traditional file systems also offer facilities to truncate, append to, create, move, delete and in-place modify files. They do not offer facilities to prepend to or truncate from the beginning of a file, let alone arbitrary insertion into or deletion from a file. The operations provided are highly asymmetric and lack the generality to be useful in unexpected contexts. For example, interprocess pipes in Unix have to be implemented outside of the file system because the pipes concept does not offer truncation from the beginning of files.
Reply With Quote

Types of file systems

File system types can be classified into disk file systems, network file systems and special purpose file systems.

Disk file systems

A disk file system is a file system designed for the storage of files on a data storage device, most commonly a disk drive, which might be directly or indirectly connected to the computer. Examples of disk file systems include FAT [FAT12, FAT16, FAT32, exFAT], NTFS, HFS and HFS+, HPFS, ext2, ext3, ext4, ISO 9660, ODS-5, ZFS and UDF. Some disk file systems are journaling file systems or versioning file systems.

Flash file systems

A flash file system is a file system designed for storing files on flash memory devices. These are becoming more prevalent as the number of mobile devices are increasing, and the capacity of flash memories increase.

While a disk file system can be used on a flash device, this is suboptimal for several reasons:

Erasing blocks: Flash memory blocks have to be explicitly erased before they can be rewritten. The time taken to erase blocks can be significant, thus it is beneficial to erase unused blocks while the device is idle.
Random access: Disk file systems are optimized to avoid disk seeks whenever possible, due to the high cost of seeking. Flash memory devices impose no seek latency.
Wear levelling: Flash memory devices tend to wear out when a single block is repeatedly overwritten; flash file systems are designed to spread out writes evenly.

Database file systems

A new concept for file management is the concept of a database-based file system. Instead of, or in addition to, hierarchical structured management, files are identified by their characteristics, like type of file, topic, author, or similar metadata.

Transactional file systems

Each disk operation may involve changes to a number of different files and disk structures. In many cases, these changes are related, meaning that it is important that they all be executed at the same time. Take for example a bank sending another bank some money electronically. The bank's computer will "send" the transfer instruction to the other bank and also update its own records to indicate the transfer has occurred. If for some reason the computer crashes before it has had a chance to update its own records, then on reset, there will be no record of the transfer but the bank will be missing some money.

Transaction processing introduces the guarantee that at any point while it is running, a transaction can either be finished completely or reverted completely (though not necessarily both at any given point). This means that if there is a crash or power failure, after recovery, the stored state will be consistent.

This type of file system is designed to be fault tolerant, but may incur additional overhead to do so.

Journaling file systems are one technique used to introduce transaction-level consistency to filesystem structures.

Network file systems

A network file system is a file system that acts as a client for a remote file access protocol, providing access to files on a server. Examples of network file systems include clients for the NFS, AFS, SMB protocols, and file-system-like clients for FTP and WebDAV.

Special purpose file systems

A special purpose file system is basically any file system that is not a disk file system or network file system. This includes systems where the files are arranged dynamically by software, intended for such purposes as communication between computer processes or temporary file space.

Special purpose file systems are most commonly used by file-centric operating systems such as Unix. Examples include the procfs (/proc) file system used by some Unix variants, which grants access to information about processes and other operating system features.

Deep space science exploration craft, like Voyager I & II used digital tape-based special file systems. Most modern space exploration craft like Cassini-Huygens used Real-time operating system file systems or RTOS influenced file systems. The Mars Rovers are one such example of an RTOS file system, important in this case because they are implemented in flash memory.

Crash counting is a feature of a file system designed as an alternative to journaling. It is claimed that it maintains consistency across crashes without the code complexity of implementing journaling.

File systems and operating systems

Most operating systems provide a file system, as a file system is an integral part of any modern operating system. Early microcomputer operating systems' only real task was file management — a fact reflected in their names. Some early operating systems had a separate component for handling file systems which was called a disk operating system. On some microcomputers, the disk operating system was loaded separately from the rest of the operating system. On early operating systems, there was usually support for only one, native, unnamed file system; for example, CP/M supports only its own file system, which might be called "CP/M file system" if needed, but which didn't bear any official name at all.

Because of this, there needs to be an interface provided by the operating system software between the user and the file system. This interface can be textual or graphical. If graphical, the metaphor of the folder, containing documents, other files, and nested folders is often used.

Flat file systems

In a flat file system, there are no subdirectories—everything is stored at the same (root) level on the media, be it a hard disk, floppy disk, etc. While simple, this system rapidly becomes inefficient as the number of files grows, and makes it difficult for users to organize data into related groups.

Like many small systems before it, the original Apple Macintosh featured a flat file system, called Macintosh File System. Its version of Mac OS was unusual in that the file management software (Macintosh Finder) created the illusion of a partially hierarchical filing system on top of MFS. This structure meant that every file on a disk had to have a unique name, even if it appeared to be in a separate folder. MFS was quickly replaced with Hierarchical File System, which supported real directories.

A recent addition to the flat file system family is Amazon's S3, a remote storage service, which is intentionally simplistic to allow users the ability to customize how their data is stored. The only constructs are buckets and objects. Advance file management is allowed by being able to use nearly any character including '/' in the objects name, and the ability to select subsets of the bucket's content based on identical prefixes.

File systems under Unix-like operating systems

Unix-like operating systems create a virtual file system, which makes all the files on all the devices appear to exist in a single hierarchy. This means, in those systems, there is one root directory, and every file existing on the system is located under it somewhere. Unix-like systems can use a RAM disk or network shared resource as its root directory.

Unix-like systems assign a device name to each device, but this is not how the files on that device are accessed. Instead, to gain access to files on another device, the operating system must first be informed where in the directory tree those files should appear. This process is called mounting a file system. For example, to access the files on a CD-ROM, one must tell the operating system "Take the file system from this CD-ROM and make it appear under such-and-such directory". The directory given to the operating system is called the mount point - it might, for example, be /media. The /media directory exists on many Unix systems and is intended specifically for use as a mount point for removable media such as CDs, DVDs and like floppy disks. It may be empty, or it may contain subdirectories for mounting individual devices. Generally, only the administrator or root may authorize the mounting of file systems.

File systems under Linux

Linux supports many different file systems, but common choices for the system disk include the ext family such as ext2 and ext3, XFS, JFS and ReiserFS.

File systems under Solaris

The Sun Microsystems Solaris operating system in earlier releases defaulted to UFS for bootable and supplementary file systems. Solaris defaulted to, supported, and extended UFS.
Support for other file systems and significant enhancements were added over time, including Veritas Software Corp. VxFS, Sun Microsystems QFS, Sun Microsystems UFS, and Sun Microsystems ZFS.

Kernel extensions were added to Solaris to allow for bootable Veritas VxFS operation. Logging or Journaling was added to UFS in Sun's Solaris 7. Releases of Solaris 10, Solaris Express, OpenSolaris, and other open source variants of the Solaris operating system later supported bootable ZFS.

Logical Volume Management allows for spanning a file system across multiple devices for the purpose of adding redundancy, capacity, and/or throughput. Legacy environments in Solaris may use Solaris Volume Manager Multiple operating systems may use Veritas Volume Manager. Modern Solaris based operating systems eclipse the need for Volume Management through leveraging virtual storage pools in ZFS.

File systems under Mac OS X

Mac OS X uses a file system that it inherited from classic Mac OS called HFS Plus. HFS Plus is a metadata-rich and case preserving file system. Due to the Unix roots of Mac OS X, Unix permissions were added to HFS Plus. Later versions of HFS Plus added journaling to prevent corruption of the file system structure and introduced a number of optimizations to the allocation algorithms in an attempt to defragment files automatically without requiring an external defragmenter.

Filenames can be up to 255 characters. HFS Plus uses Unicode to store filenames. On Mac OS X, the filetype can come from the type code, stored in file's metadata, or the filename.

HFS Plus has three kinds of links: Unix-style hard links, Unix-style symbolic links and aliases. Aliases are designed to maintain a link to their original file even if they are moved or renamed; they are not interpreted by the file system itself, but by the File Manager code in userland.

Mac OS X also supports the UFS file system, derived from the BSD Unix Fast File System via NeXTSTEP. However, as of Mac OS X 10.5, Mac OS X can no longer be installed on a UFS volume, nor can a pre-Leopard system installed on a UFS volume be upgraded to Leopard.

File systems under Microsoft Windows

Windows makes use of the FAT and NTFS file systems.

The File Allocation Table (FAT) filing system, supported by all versions of Microsoft Windows, was an evolution of that used in Microsoft's earlier operating system. FAT ultimately traces its roots back to the short-lived M-DOS project and Standalone disk BASIC before it. Over the years various features have been added to it, inspired by similar features found on file systems used by operating systems such as Unix.
Older versions of the FAT file system (FAT12 and FAT16) had file name length limits, a limit on the number of entries in the root directory of the file system and had restrictions on the maximum size of FAT-formatted disks or partitions. Specifically, FAT12 and FAT16 had a limit of 8 characters for the file name, and 3 characters for the extension (such as .exe). This is commonly referred to as the 8.3 filename limit. VFAT, which was an extension to FAT12 and FAT16 introduced in Windows NT 3.5 and subsequently included in Windows 95, allowed long file names (LFN). FAT32 also addressed many of the limits in FAT12 and FAT16, but remains limited compared to NTFS.

NTFS, introduced with the Windows NT operating system, allowed ACL-based permission control. Hard links, multiple file streams, attribute indexing, quota tracking, compression and mount-points for other file systems are also supported, though not all these features are well-documented.

Unlike many other operating systems, Windows uses a drive letter abstraction at the user level to distinguish one disk or partition from another. For example, the path C:\WINDOWS represents a directory WINDOWS on the partition represented by the letter C. The C drive is most commonly used for the primary hard disk partition, on which Windows is usually installed and from which it boots. This "tradition" has become so firmly ingrained that bugs came about in older versions of Windows which made assumptions that the drive that the operating system was installed on was C. The tradition of using "C" for the drive letter can be traced to MS-DOS, where the letters A and B were reserved for up to two floppy disk drives. Network drives may also be mapped to drive letters.
Data retrieval process

The operating system calls on the IFS manager. The IFS calls on the correct FSD in order to open the selected file from a choice of four FSDs that work with different storage systems—NTFS, VFAT, CDFS, and Network. The FSD gets the location on the disk for the first cluster of the file from the FAT, FAT32, VFAT, or, in the case of Windows NT based, the MFT. In short, the whole point of the FAT, FAT32, VFAT, or MFT is to map out all the files on the disk and record where they are located.


Source : techarena

May 3, 2008

Database Terms

Database Terms

  • Aggregation - Objects that are composed of other objects.
  • Concurrency - Databases must ensure that data is checked when concurrent access is allowed. Concurrent access means more than one application or thread may be reading or updating the same data at the same time.
  • Database Garbage Collection - Garbage collection is the process of destroying objects that are no longer referenced, and freeing the resources those objects used. In Java there is a background process that performs garbage collection. Requires bi-directional object relationships. Determines if the database performs garbage collection on objects that are no longer referenced by the database. This keeps external programs from having to track the use of object pointers.
  • DBMS - Database management system.
  • Distributed Architecture - Object are sharing in a distributed environment or the entire database may be replicated on multiple computers.
  • DML - Data Manipulation Language separate from programming languages (for RDBMS) and used as a means of getting and storing data in the database.
  • Encapsulation - Data with method storage. Not all databases support the methods but rely upon the classes defined in the schema to reconstruct the object with its methods.
  • Fault tolerance - Features that provide for fault tolerence in the event of a hardware of software failure. Normally transaction processing provides software fault tolerance. Data replication to other servers on the network supports hardware fault tolerance.
  • Heterogeneous environment - Cross Platform support - The database may be able to run on various builds of computers and with various operating systems.
  • Inheritance - Objects inherit attributes from parent objects.
  • JDBC - An application program interface (API). Calls are used to execute SQL operations.
  • normalized - Elimination of redundancy in databases so that all columns depend on a primary key.
  • Notification - Notification may be active or passive. A passive system can minimally determine if an object has changed state. An active system may provide for an application to be informed when an object is modified.
  • Object relationships - Object relationships define association with other objects, and whether objects can detect each other in one direction or two directions. Two way object relationships may allow for garbage collection.
  • ODBMS - Object Database management system.
  • OQL - Object Query Language, is a data manipulation language for Object Databases although many object databases do not support it. They rely on object class extensions or interfaces for their support.
  • Persistence - Databases provide persistance which for object databases means object can be stored between database runs.
  • RDBMS - Relational database management system.
  • Schema - The data structure of the database.
  • SQL - Structured Query Language is a standard language for communication with a relational database management system (RDBMS). Structured Query Language, is a data manipulation language which is a standard for getting and storing data in an RDBMS.
  • SQLJ - Supports structured query language (SQL) calls for Java. It consists of a language allowing SQL statements to be embedded in it, a translator, and a runtime model.
  • Transaction processing - Some databases may have some form of transaction processing which may support concurrency. Transaction processing will ensure that the entire transaction is made or none of it is made. Transactions support concurrency and data recovery. A data failure will cause a rollback of data

ORDBMS Definition

ORDBMS Definition

An object relational database is also called an object relational database management system (ORDBMS). This system simply puts an object oriented front end on a relational database (RDBMS). When applications interface to this type of database, it will normally interface as though the data is stored as objects. However the system will convert the object information into data tables with rows and colums and handle the data the same as a relational database. Likewise, when the data is retrieved, it must be reassembled from simple data into complex objects.

Performance Constraints

Because the ORDBMS converts data between an object oriented format and RDBMS format, speed performance of the database is degraded substantially. This is due to the additional conversion work the database must do.

ORDBMS Benefits

The main benefit to this type of database lies in the fact that the software to convert the object data between a RDBMS format and object database format is provided. Therefore it is not necessary for programmers to write code to convert between the two formats and database access is easy from an object oriented computer language.

Object Oriented Database Standards

Object Oriented Database Standards

There are several object oriented standards and groups that oversee them.

Groups

  • Object Management Group (OMG) - Develops standards to help make object applications to be portable and communicate between each other (interoperability). They have developed the Component Object Request Broker Architecture (CORBA) standard along with object and OODBMS interfaces.
  • Object Database Management Group (ODMG) - Created to define standard interfaces for object databases. The interfaces should allow the databases and applications that use them be portable and communicate between each other.

Standards

  • CASE Data Interchange Format (CDIF) - Defines standards for tools to use so tyey may be used with various applications such as database servers, application servers, and other tools.
  • Portable Common Tool Environment (PCTE) Standard.
  • PDES/STEP - Exchange format standard for product model data. Also called the object interface format.
  • Object Query Language (OQL)
  • Object Definition Language (ODL) - Extension of OMGs CORBA standard.

Object Database Use and Features

Object Database Use and Features

Databases provide persistence which for object databases means that objects can be stored between database runs.

Features

The following list of features are capabilities that object databases may support. Object database features include:

  • Support of the object oriented language you want to use.
  • Support of Object Oriented Concepts.
    • Aggregation - Objects that are composed of other objects.
    • Encapsulation - Data with method storage. Not all databases support the methods but rely upon the classes defined in the schema to reconstruct the object with its methods.
    • Inheritance - Objects inherit attributes from parent objects.
    • Polymorphism - Allows two methods to use the same name but have different behavior. Methods for one object can be defined, then the operation specification can be shared with other objects.
  • Distributed Architecture - Object are sharing in a distributed environment or the entire database may be replicated on multiple computers.
  • Heterogeneous environment - Cross Platform support - The database may be able to run on various builds of computers and with various operating systems.
  • Transaction processing - Some databases may have some form of transaction processing which may support concurrency. Transaction processing will ensure that the entire transaction is made or none of it is made. Transactions support concurrency and data recovery. A data failure will cause a rollback of data.
  • Concurrency - Databases must ensure that data is checked when concurrent access is allowed. Concurrent access means more than one application or thread may be reading or updating the same data at the same time. This may also be called two phase commit where two processes may work on the same object at the same time. This may use data locking for reads or writes. Some methods of concurrency control include:
    • Pessimistic control - Whan one or more processes are reading, updated to the data cannot be made.
    • Multiread - Updates are not blocked. Data must be consistant when the transaction was begun. In other words, if the read was done, and the data was changed by another process before the data is saved, the transaction is not valid until the data is read again.
  • Object relationships - Object relationships define association with other objects, and whether objects can detect each other in one direction or two directions. Two way object relationships may allow for garbage collection. The best option is two way relationships.
  • Database Garbage Collection - Requires bi-directional object relationships. Determines if the database performs garbage collection on objects that are no longer referenced by the database. This keeps external programs from having to track the use of object pointers.
  • Relationship cardinality - Supported relationships may include any combination of:
    • One to one.
    • One to many.
    • Many to one.
    • Many to many.

The database should support all these.

  • Transparent persistence - Consists of direct data manipulation using object oriented language. Many times a persistent capable class or persistent interface is used to implement persistence. This may be considered by some vendors to be transparency. If an interface is used, an intermediate interface may be used to help insulate calls from the particular database, thereby allowing the customer to more easily change database vendors later.
  • Database interface methods - These may include SQL, OQL, and some application programming interface (API). See the section under "Communication Support", below.
  • Database Integrity - There are two types:
    • Structural database integrity ensures database contents are consistent with the database schema. Referential Integrity requires bi-directional object relationships to ensure objects do not contain references to deleted objects.
    • Logical integrity - The logical properties of the data are correct. The data has the correct values consistently and concurrent access does not cause incorrect values to be set.
  • Object Versioning - A single object represented by multiple versions. Two types are:
    • Linear - Prior versions of the object are saved as the object is changed.
    • Branch - Multiple users may update the object concurrently.
  • Notification - Notification may be active or passive. A passive system can minimally determine if an object has changed state. An active system may provide for an application to be informed when an object is modified.
  • Indexing - Additional indexing may be provided to enhance data retrieval efficiency. Hashing and b-trees may be used.
  • Security - Data storage and/or transmission encryption may be supported by some databases. Also different authentication methods and levels for access to the database may be provided by various products.
  • Archiving and data recovery
  • Fault tolerance - Features that provide for fault tolerence in the event of a hardware of software failure. Normally transaction processing provides software fault tolerance. Data replication to other servers on the network supports hardware fault tolerance.
  • Data access - Access is normally done using an iterator to access the data as though objects are collections. This way the objects are not required to be loaded into memory before the desired object is obtained.
  • Sorting - All objects of a given class or parent class may be obtained.
  • Tools that can be used with the database.
  • Amount of storage.

Method storage- The code that runs in objects and gives them behavior is stored in the database.

Considerations

  • What programming languages does the database support?
  • Object relationships - Are they bi-directional?
  • Work Group Support - Sharing databases and locking.
  • Schema Evolution - How do you tell the database about schema changes? This includes changes to the definition of a class such as attributes or behavior, changes to inheritance, adding, deleting, or renaming a class. Do classes need to be backward compatible?
  • How do databases search using polymorphism? Can it give all cars objects that are made by a specific manufacturer?
  • How are the database APIs used? Is the database transparent to the applications?
  • Tools - Tools are important for product development and support. The database may support some tools or integrate with some. Tools that should be a concern include development tools, testing tools, debugging tools, data modeling tools, and data maintenance tools.
  • Object Models - The object modeling to be used and whether the object modeling tools integrate with the database (or whether they should) should be considered.
  • Does the database store object methods or rebuild the methods from classes when required? If it does store methods, methods can be executed in database processes without storing the method or recreating the method in the application memory. Non object oriented programs may be able to access the output of the stored methods.

Communications Support

Object databases will use one or more of the following methods to exchange data between applications and the database.

  • OQL - The standard language for object database communication is object query language (OQL). Some object databases support it and others do not.
  • SQL - The standard language for relational database communication is structured query language (SQL). Some object databases support it and others do not. This is provided to help prospective customers migrate current applications from RDBMS to ODBMS. The object databases use SQL by considering a row an object, and each unit in a column to be an attribute of an object. The table is a collection of objects. The table joins and keys are used to create object relationships.
  • Application Programming Interface (API) - Some vendors provide additional classes or programming interfaces that are used to access the database. It consists of direct data manipulation using object oriented language.

The advantage to using a standard interface such as SQL or OQL is that the application is not tied to one specific database. The advantage to using an API is that the access may be faster and possibly even transparent to the application. The application may not even know it is running methods or using data on a database. The API is a mixed bag since it gives some performance advantages and perhaps a little less flexibility. This loss of flexibility may be mitigated by providing a standard interface between the application and the particular database's API.

Main Comparison Points

This table is an evaluation example and may not reflect all languages and platforms supported along with possible other features.

Feature

Objectivity

Vendor 2

Vendor 3

Supported Languages

Java



Supported Systems

Linux, NT



Object oriented concepts

yes



Distributed

yes



Heterogeneous

yes



Transaction Processing

yes



Concurrency

Multiple readers, single writer



Object relationships

Two way



Garbage collection

yes



Relationship cardinality

all 4



Transparent persistence

yes/sort of



Interface Methods

Persistent capable objects/interfaces, SQL



Database integrity (Logical/Structural)

Logical and structural



Versioning (Linear or Branch)

Linear



Notification (Active or Passive)

Passive



Additional Indexing

Minimal



Security - Data encryption

yes



Security - Secure authentication

yes



Security levels available?

yes



Data storage and recovery

yes



Fault tolerance

yes



How Objects are Stored

  • Object Identifiers (OIDs) are used.
  • Relationships are constructed and tracked.

Object Oriented Databases

Object Oriented Databases

Object oriented databases are also called Object Database Management Systems (ODBMS). Object databases store objects rather than data such as integers, strings or real numbers. Objects are used in object oriented languages such as Smalltalk, C++, Java, and others. Objects basically consist of the following:

  • Attributes - Attributes are data which defines the characteristics of an object. This data may be simple such as integers, strings, and real numbers or it may be a reference to a complex object.
  • Methods - Methods define the behavior of an object and are what was formally called procedures or functions.

Therefore objects contain both executable code and data. There are other characteristics of objects such as whether methods or data can be accessed from outside the object. We don't consider this here, to keep the definition simple and to apply it to what an object database is. One other term worth mentioning is classes. Classes are used in object oriented programming to define the data and methods the object will contain. The class is like a template to the object. The class does not itself contain data or methods but defines the data and methods contained in the object. The class is used to create (instantiate) the object. Classes may be used in object databases to recreate parts of the object that may not actually be stored in the database. Methods may not be stored in the database and may be recreated by using a class.

Comparison to Relational Databases

Relational databases store data in tables that are two dimensional. The tables have rows and columns. Relational database tables are "normalized" so data is not repeated more often than necessary. All table columns depend on a primary key (a unique value in the column) to identify the column. Once the specific column is identified, data from one or more rows associated with that column may be obtained or changed.

To put objects into relational databases, they must be described in terms of simple string, integer, or real number data. For instance in the case of an airplane. The wing may be placed in one table with rows and columns describing its dimensions and characteristics. The fusalage may be in another table, the propeller in another table, tires, and so on.

Breaking complex information out into simple data takes time and is labor intensive. Code must be written to accomplish this task.

Object Persistence

With traditional databases, data manipulated by the application is transient and data in the database is persisted (Stored on a permanent storage device). In object databases, the application can manipulate both transient and persisted data.

When to Use Object Databases

Object databases should be used when there is complex data and/or complex data relationships. This includes a many to many object relationship. Object databases should not be used when there would be few join tables and there are large volumes of simple transactional data.

Object databases work well with:

  • CAS Applications (CASE-computer aided software engineering, CAD-computer aided design, CAM-computer aided manufacture)
  • Multimedia Applications
  • Object projects that change over time.
  • Commerce

Object Database Advantages over RDBMS

  • Objects don't require assembly and disassembly saving coding time and execution time to assemble or disassemble objects.
  • Reduced paging
  • Easier navigation
  • Better concurrency control - A hierarchy of objects may be locked.
  • Data model is based on the real world.
  • Works well for distributed architectures.
  • Less code required when applications are object oriented.

Object Database Disadvantages compared to RDBMS

  • Lower efficiency when data is simple and relationships are simple.
  • Relational tables are simpler.
  • Late binding may slow access speed.
  • More user tools exist for RDBMS.
  • Standards for RDBMS are more stable.
  • Support for RDBMS is more certain and change is less likely to be required.

ODBMS Standards

  • Object Data Management Group
  • Object Database Standard ODM6.2.0
  • Object Query Language
  • OQL support of SQL92

How Data is Stored

Two basic methods are used to store objects by different database vendors.

  • Each object has a unique ID and is defined as a subclass of a base class, using inheritance to determine attributes.
  • Virtual memory mapping is used for object storage and management.

Data transfers are either done on a per object basis or on a per page (normally 4K) basis.

RDBMS Definition

RDBMS Definition

Relational databases store data in tables (relations) that are two dimensional. The tables have rows (records or objects) and columns (fields or attributes). Data items at an intersection of a row and a column are called a cell and consist of attribute values. Data stored is simple data such as integers, real numbers or string values. Multiple values may not be stored in one cell. Relational database tables are "normalized" so data is not repeated more often than necessary. All table columns depend on a primary key (a unique value in the column) to identify the column. Once the specific column is identified, data from one or more rows associated with that column may be obtained or changed.

Relational databases are sets of tables. One table file is not a relational database. A relational database server is not the same as a relational database. A relational database can be a file with sets of tables. The relational database server includes the ability to service requestesto get or change data from remote clients.

Relational database servers use Structured Query Language (SQL), as a data manipulation language to interface between itself and the clients. SQL is the standard for getting and storing data in an RDBMS. For information about SQL, see the "Beginner's SQL Guide".

Relational database servers provide:

  • Data Management
  • Transaction processing
  • Data integrity - Provides for multiple access at the same time (concurrency) between multiple processes/users. This is done so data is not displayed nor saved in a fashion where one change is lost. Various locking mechanisms are used to support this.
  • Data backup and recovery.
  • Data security - Provides for user authentication, and levels of data access privileges.

Primary Keys

In relational databases, the data is not arranged in any particular order in tables. The data in tables requires keys for identification of rows. Each table has rows and columns. Sets of values in a row may describe a particular item such as customers. Consider the following table:

Customers

Customer ID

Name

Street Address

City

Zip Code

Phone

e-mail

112304

John Brown

123 Straight Lane

Paradise

MI

49555

johnb@heaven.net

134056

Tim Smith

321 Curved Ave

Hollywood

MD

10255

tims@eastnet.net

234902

Sue Jones

213 Hill Road

Knoxville

TN

23555

suej@nicenet.net

092387

George Abernathy

555 North Road

Smalltown

FL

67555

georgea@mynet.net

187462

Tom Jenkins

666 South Lane

Cold

MN

55555

tomj@networks.com

108976

Kim Adams

456 Peachtree Lane

Atlanta

GA

54555

kima@bestnet.net

059385

Chris Christopher

859 East Road

Nicetown

MS

45555

chrisc@bignet.com

140763

Jim Thompson

307 West Ave

Bigcity

ND

65555

jimt@fastnet.net

Columns include Customer ID, Name, Street Address, City and so forth. Each column has a different set of information about all customers such as their phone number. This information could be considered customer characteristics or attributes. Each row has all information about one particular customer. Each row must have a unique means of identification. We could have used the customer name to identify the customer, but it is possible that more than one customer may have the same name. Therefore we are using unique customer identifiers. The Customer ID is the key used to identify individual customers. The column to be used as the key must be identified to the relational database. This is called the primary key.

Foreign Keys

Let's say we're running a store with the customers listed above. The first table below lists items for sale. The second table lists orders that have been placed in the store, and the third table lists items ordered for an individual order. If we want to print out what items customer 234902 ordered we do the following:

  1. Go to the orders table and get all orders customer 234902 ordered.
  2. For each order found (1 order #190389575). The order number is considered to be a foreign key to another table and column which is to table "Individual Order 190389575" and Item ID.
  3. Go to the "Individual Order 190389575" table and get each item ID that was ordered. Each item ID is a foreign key to the "Store Items" table and Item ID column.
  4. Go to the store items table, find each item ID and print out each item attribute such as price, description, and quantity purchased (from the individual order table).

Store Items

Item ID

Description

Price

000234567

Small Flower Pot

2.36

000018901

Potting Soil

5.97

001034654

Geranium Seed

1.50

Orders

Order ID

Customer ID

Subtotal

Tax

Total

190389575

234902

16.16

0.82

16.98

109748230

187462

2.37

0.12

2.49

208949023

059385

104.23

5.22

109.45

103792034

187462

40.00

2.00

42.00

123048938

134056

10.00

0.50

10.50

Individual Order 190389575

Item ID

Quantity

000234567

2

000018901

2

001034654

1

Therefore foreign keys are used to access data in related tables.

Server Types

Server Types

There are several kinds of databases which include:

  • Flat file databases
  • Relational databases
  • Object databases
  • Object relational databases

Flat files are simply files with a table of information which may be seperated by delimeters such as commas, colons, or semi-colons. Relational databases consist of several related tables of simple data. The tables are composed of rows and columns. Object databases store data in an object form rather than in tables. They store attributes and class information, but sometimes they also store and methods (behavior) in the database. Object relational databases are relational databases with data stored in tables, but they have a front end that converts objects to data and data to objects, making it seem to the application that objects are being stored.

Database servers include both a server program that serves remote clients and manages the database. They may use some means of standard communication between client and server to allow management of the data such as structured query language (SQL) for relational database servers.

The most popular databases today are Relational Database Management Systems (RDBMS). However, object database servers may someday overtake the relational database servers.

Relational Databases and Objects

If relational databases are used to store objects, the object must first be disassembled into parts, normalized, and placed in tables. This can take some time to do, and be a labor intensive process required for writing the code. To use the object, it must be reassembled.

Many relational databases are run on one single server and do not use a distributed architecture.

Object Database Servers

Object database servers may use an Object Query Language (OQL) as a standard language for communication. They may use an application programming interface (API) to allow the application to control the data or they may use both the API and OQL.

Current and Future Trends

Relational databases are still the most popular database in use today. There is good reason for this. They are easy to use and are normally efficient.

However as programming has changed, tools related to those changes must also change. Object oriented programming is becomming much more popular and as that occurs a more practical tool for long term storage of data is in demand. This tool must interface easily to the object oriented language in question. It must also be a standard tool so users are not tied to specific vendors amd should have a standard way of exchanging information between applications and the database. OQL was developed for this purpose, but it does not appear to be widely supported yet by object oriented database vendors.

Although object databases were first written many years ago, since they have not yet become popular, it appears that the market is not stable. There are several object oriented database vendors, and it is difficult to tell who will be in the market for the long haul. Therefore, I believe the purchase of an object oriented database is somewhat of a risk. This risk may be somewhat mitigated by the fact that programs can be written in object oriented language to isolate the programs from specific object database products.

If the benefits of the object oriented database are large enough for the particular application they are used for and the specific organization considering them, the risks are likely to be worth taking. However, if the benefits are marginal, it may be worth waiting another year or two for more market stability and uniformity

Popular Posts