Wednesday, May 18, 2011

Vision for the School of Information Sciences

Statement

Vision

To foster a world where digital information is used optimally to advance the information age and improve the human condition.

Mission

To conduct research that advances our understanding of information and knowledge in all forms and across all operations; and to produce professionals who can guide individuals, organizations and society in the use of information to the betterment of humanity.

Strategic Goal
To advance our research and scholarship, so as to become world class, in the following areas of information science: information assurance, archival science, wireless communication, school librarianship, distributed systems, cyber-scholarship. (The areas selected are somewhat arbitrary and should be modified to reflect faculty consensus.)

Operational Goals

  • To provide the tools and infrastructure necessary for the faculty and PhD students to implement preliminary research studies in areas of strategic interest and a proposal preparation and review process that enables these studies to be successfully implemented as externally funded research projects.
  • To provide the tools, training, and infrastructure required to mount educational programs that prepare individuals for roles as information professionals and scientists.
  • To represent the profession and science in public, professional and scientific forums in such a way as to advance the vision, mission, and strategic goal of the school.
Background
In order to understand the vision and mission for the School of Information Sciences put forward above, it is useful to reflect on what we mean by information, profession, and science. The School has made a conscious decision to say that it the home of more than one science of information. At the same time, we have been lax in defining exactly what those sciences are – or defining the combinations of professions and sciences that are involved. What follows makes no judgments about the particular sciences, but it does endeavor to provide some definitional perspective on information, profession, and science.

What is Information?
There are both informal and formal definitions of information. Conversationally, a message communicated from one person to another is said to contain information if the message content was not already known to the receiver. If the message was known to the receiver, it contains no information or at very least no “new” information. This raises the question of whether recorded data can be defined as containing information apart from some person. We commonly say that containers like dictionaries, encyclopedias, or books “contain a lot of information”. It might or might not be information to you. These informal definitions suggest two things. First, information is intimately tied to the human experience and information is in some forms personal. What is information for you may not be information to me. This is a separate matter from the veracity of an assertion or message. Second, we recognize that information may be collected and organized in a variety of forms. We are able to form impersonal judgments about these collections as being rich or poor.

Is it possible to "inform" an ant, building, computer, or automobile? I am not quite sure. While information is tightly bound to the human experience, some would argue that ant scent trails contain information, and that computers produce information displays -- sometimes regardless of whether they are used by humans. It is clear to me that computer programs make more and more use of information to make decisions about data collected from the real world. It seems likely that an informal definition of information will likely grow to include many types. The argument here is simply that information is an artifact of human efforts and is closely tied to the human experience.

In terms of more formal definitions of information, Claude Shannon’s work on cryptography at Bell Labs during World War II led to a mathematical definition of information. Put in the simple terms Shannon suggested that the greater the uncertainty (entropy) in a message, the more information required to transmit the outcome. Again, at the risk of oversimplification, communicating the outcome of some event where there are a dozen possible outcomes required more information than an event with only two possible outcomes. Working for Bell Labs, Shannon was interested in how much bandwidth was needed to communicate a message, or h
ow much space was needed to store a message. With the advent of digital computer using binary units to store a message, it became useful to think about how many binary digits would be required to store a message. Shannon proposed a measure of information was "I=log (1/probability of the message)". We can thus say that if information is to be stored as an array of binary digits and the probability of the message is .5, we would need log2 (1/.5) bits to store the message. The base 2 log of (1/.5) (i.e. 2) is 1. When the probability of a message is 50/50, I can record the message as a 0 or 1 using only one binary digit. If the probability of a given message is .00390625 I would need log2 1/.00390625 or 8 bits. In this case, the “magic” number of .00390625 is 1/256. What Shannon is saying is that if I wish to be able to have a message that can be any one of 256 different symbols, I would need 8 bits to represent it. A rich analysis of information is possible based on a measure of information as the probability of a message. It opens the doors to computation, transmission, encryption, compression, correction of messages stored in digital form. It provides a simple yet rich mathematical theory that allows us to do all sorts of things, and it is generally consistent with the informal definition of information.

Given these simple definitions of information, numerous additional questions emerge. For example, what it is that makes something information to one but not another? Messages contain information when the message is about something not already known to the receiver. Again, informally speaking, what I know, the knowledge I have, acts as the mediator of whether a message contains information. This leads us to ask if knowledge is different from a store of information. What we know may be partitioned into domains. We may know a lot about chemistry, a little about mathematics, and nothing about anthropology. It would seem safe to suggest that knowledge is generally thought to be more organized and structured than information.

This discussion began with the notion that we received information in messages. Is that the only source of information? It would seem not to be the case. Humans are able to derive information from signals processed from the environment. The smell of smoke might lead us to conclude that a fire is burning. The sound of footsteps might lead us to conclude that a person or other creature is nearby. Signals, when informed by some mechanism that allows them to be interpreted, may contain information. There is much to be discussed about how signals are transformed to a pattern of data and then to information. It has a lot to do with the knowledge we bring to bear on the signals. Signals that make no sense – that form no pattern – are commonly referred to as noise.

Information is an artifact of human effort. Frequently, messages are used to transfer information between individuals. Humans may glean information from signals gathered from the world. A message in a symbolic form may be aggregated in an information store. One measure of the information in a message is the log or the inverse of the probability of the message. Information may be mediated, both in production and acquisition by the knowledge state of a recipient. There are evident properties of knowledge that require clearer definition such as validity, applicability, coherence, etc. Similarly, there are evident properties of information that must be clearly defined, such as its value, its clarity, its validity, etc.

Profession and Science
The academic disciplines, humanities and sciences, and the professions, medicine and engineering to name two, are sometimes viewed as being the same. While there are similarities, there are also important differences. Where the disciplines are an expression of human imagination, professions are a response to human need. While both have a systematic method and theory, in a discipline, knowledge, theory, and process are mastered to develop a more systematic or comprehensive syntactical and conceptual structure for the discipline — the purpose of the discipline is the pursuit of knowledge. Professions, on the other hand, use theory and knowledge to serve — to respond to vital needs of individuals. Holzner and Marx (1979) put it this way:
Professionals differ from scientists in their reliance on applied rather than formal theoretical knowledge. Thus, practicing professionals tend to use general principles to deal with concrete problems in order to test, elaborate, or arrive at general principles. Since professionals must focus on the practical solution of concrete problems, they are characteristically more concerned with applying knowledge than with creating or contributing to it. (Holzner, B., and J. Marx, 197 Knowledge Application: The Knowledge System in Society. Boston: Allyn and Bacon. p. 407)
King and Brownell (King, A.R. and Brownell, J.A., The Curriculum and the Disciplines of Knowledge: A Theory of Curriculum Practice. New York: John Wiley and Sons, 1966) discuss the characteristics of a discipline. These include both fundamental characteristics – a domain of inquiry, a mode of inquiry, a conceptual structure – and more descriptive characteristics – a community of persons with a literature, a tradition in an instructive community and a specialized language. The three characteristics identified as fundamental are worthy of a more careful exposition.

It is pretty clear that the domain of physics is the physical universe, the domain of biology is living organisms, the domain of literature is writings, etc. At gross levels, these domains are difficult to constrain, but as we talk about astrophysics, or vertebrate biology, the domains seem to become a little more sharply defined. Sometimes, they get fuzzier -- e.g. molecular biology, or social psychology. What is the domain of information science? One less than satisfying answer is everything is information and the domain of information science includes all the other disciplines. Some suggest that we don't need to define or circumscribe the definition of information to have a science of it. Lacking a reasonable definition of living organisms would make it very difficult to define what biology is about. We don’t need to be perfect – every domain has outliers. While vertebrates and plants are clearly living organisms, there are surely fringe entities that lack one or more of the attributes we use to define living organisms. When it comes to information, it seems we only want to argue about the fringe of the domain and ignore the 99% that is at the core. So what is the domain of information science? A broad definition of information has been provided above that may serve as a starting point.

Ok, we are coming to grips with what we want to study. What is the conceptual structure we overlay on the phenomenon? In physics, we have had a number of conceptual models of the physical world, at both microscopic and macroscopic levels. Newtonian mechanics worked for a long time. Quantum mechanics takes another view, not necessarily contradictory, but in some situations, more explanatory. The Newtonian conceptual structure provides principles like force = mass * acceleration. This concept is not a part of the domain of inquiry, it is a part of the conceptual structure that is overlaid to explain something about the domain. So, what is the conceptual structure of the science of information? Some might suggest that it is Claude Shannon's conceptualization of information as the log of the sum of the inverse of the probabilities of the components of the message. While Shannon’s formula for the measure of information has some important uses, it is way too thin – the conceptual framework that needs to be defined for a science of information is much larger. As in other disciplines, that conceptual structure will evolve and face radical points of evolution over time. There are lots of bits and pieces waiting to be pulled together and organized – coding theory, data structures, logic, entropy theory, network theory. Within these domains, there are bits of specific named theories that are emerging Shannon’s Theorems, Nyquist’s Theorem, Zipf’s Law, Metcalfe’s Law, Moore’s Law, etc.

What is the method of inquiry? At the current time, there are a variety of different methods employed by people who call themselves information scientists. Viewed from a developmental point of view, many disciplines have evolved from an early period in which the primary mode of inquiry was simple observation and classification to a more evolved mode that was more formal and which enabled assessment of the validity of the conceptual model via replicable evaluation. For most sciences, this more formal method has become some variation of the scientific method. I suspect that the maturity of the science of information is such that the current period is one of observation and classification to build a base of concepts that we may later be able to relate. Eventually the work of Shannon, Zipf, Hamming, Metcalfe Nyquist and many others will be woven into a more coherent framework.

One clue about how to proceed in building a science of information was provided by Herb Simon in his 1969 “The Sciences of the Artificial”. In this book, he discussed the differences between natural and artificial sciences. He encouraged the development of a new paradigm for conducting research in the “design sciences.” He suggests that the sciences of artifacts are different than the sciences of the natural.
My dictionary defines “artificial” as “Produced by art rather than nature; not genuine or natural; affected; not pertaining to the essence of matter.” It proposes, as synonyms: affected, factitious, manufactured, pretended, sham, simulated, spurious, trumped up, unnatural. As antonyms, it lists: actual, genuine, honest, natural, real, truthful, unaffected. Our language seems to reflect man’s deep distrust of his own products. I shall not try to assess the validity of that evaluation or explore the possible psychological roots. But you will have to understand me as using “artificial” in as neutral a sense as possible, as meaning man-made as opposed to natural. (2nd edition, page 6)
And:
… hence we can set the boundaries for sciences of the artificial:
1. Artificial things are synthesized (though not always or usually with full forethought) by man.
2. Artificial things may imitate appearances in natural things while lacking, in one or many respects, the reality of the latter.
3. Artificial things can be characterized in terms of functions, goals, and adaptation.
4. Artificial things are often discussed, particularly when they are being designed, in terms of imperatives as well as descriptives.(2nd edition, page 8)
Information is an artifact of the human effort to communicate. Thus, information science is an artificial science, and not a natural science. Natural sciences endeavor to describe and explain the natural world around us where the natural world is a given. Artificial sciences endeavor to improve the design of the artifacts that we create. A science of artifacts or the artificial is a science of the things we build. In academia, there is strong pressure to do good research. Many times this is equated to descriptive and explanatory research focused on the natural world around us. We can’t make a pulsar something it is not. We simple try to explain it. In their research, engineers would be like physicists and doctors would be like biologists. Maybe we need to rethink the paradigm of our science, more focused on the matter of our science – artifacts – than on the paradigms of those who study nature.

As a simple example, if we find that we can’t communicate what we wish to via the existing symbol set, we can change the symbol set. As another example, if we find a sequenced set of symbols does not provide adequate facility to manipulate and control a document, we can define document as a directed acyclic graph of elements over that symbol set. This might allow us to do partial locking and structural analysis. We could take the example further and introduce attributes and metadata to the model to give us additional capability. This kind of design science is very different from the descriptive natural science that says a document is what it is and it is our goal to describe it in its natural form.

Michael B Spring, May 18, 2011

No comments:

Post a Comment