Showing posts with label CERIF. Show all posts
Showing posts with label CERIF. Show all posts

Friday, 9 September 2011

RSP Readiness for REF (R4R) workshop, 5th September 2011


In the first of two guest posts, Neil Stewart reflects on the RSP Readiness for REF workshop.

If you would like to contribute a guest-post to the UKCoRR blog, please contact a member of the committee.
___________________________________________________________

My name is Neil Stewart, and I'm the repository manager for the newly minted City Research Online repository, at City University London. I normally blog at City Open Access, if you want to keep an eye on developments of a repository which is still on a project footing, rather than a fully-fledged service.

The reason you find me writing here is because I was recently the recipient of two invitations: to present at the RSP Readiness for REF (R4R) workshop, held on Monday 5th September in London, and to blog about that event here at UKCORR's blog. I was happy to take up both invitations. What follows summarises some thoughts about the workshop, and on the Research Excellence Framework (REF) as it relates to repositories in general. It should be noted that the opinions below are my own, and do not necessarily reflect those of UKCORR. I'll be writing another post soon, which outlines the R4R workshop presentation which I delivered with my former LSE colleague Dave Puplett.

The event's emphasis (apart from the presentation which Dave & I delivered) can fairly be characterised as macro-level, since it discussed REF data submission in very general terms, and with specific reference to the CERIF metadata standard. CERIF is a flexible, extensible model for metadata about research-producing institutions, including (but not limited to) publications data. It is possible, using the CERIF schema, to model an institution's structure, then show how researchers and research outputs relate to that structure. This has obvious benefits for an exercise like the REF, which will (amongst other things) require submission of data on "REF-able" (i.e. high quality) publications at a department (or at least department-like "unit of assessment") level.

The morning sessions dealt with laying out the details of how CERIF and CERIF-compliant repositories could assist with REF submissions. The first session was an overview of the JISC-sponsored Readiness4REF (R4R) Project, and was delivered by Richard Gartner of Kings College London. R4R, just in the process of closing, has looked at the way in which repositories and other data management tools can provide CERIF-compliant data for REF submission purposes. Second was Keith Jeffery from euroCRIS, who gave the bigger picture as to how CERIF was developed and what it can do. This was followed by a panel discussion on the Measuring Impact Under CERIF (MICE) project, which is attempting to build in research impact data into the CERIF schema, and hence make it readily submissible for REF.

After Dave & my presentation and lunch, there were demonstrations of R4R plug-ins for the three major repository software types (Fedora, DSpace and ePrints). As an ePrints user, I was interested to see a demonstration of ePrints v. 3.3, which is to be released in the next few weeks, and contains some "CRIS-like" functionality. This is a kind of CERIF-lite approach, by which it is easy to create associations between researchers, research grants, research centres etc. to express CERIF-like linkages between them. These linkages can then be exposed in useful ways using ePrints web pages, but also exported as CERIF data to be re-used in other systems, or for a REF submission. This seems to me an interesting development, and one we may have to look at here at City.

The final session featured the obligatory break-out groups. I was assigned to a group which discussed the question: "Do you think CERIF is now a more viable option for your institution to use for its REF submission?" A variety of subjects were covered, as ever with these type of discussions. The two main points I took from it were the fact that CERIF provides the opportunity to provide an "open" citation model by modelling linkages (including positive and negative citations) between publications, outside of the "walled gardens" provided by Scopus and Web of Science; and that, for CERIF to work within my institution, there is the somewhat intractable problem of knowing to whom to speak to find out if, for example, the HR database can be plugged into the repository to transfer CERIF-formatted data between the two systems.

All in all, an interesting and timely event. Keep an eye out for my post on LSE Library's experiences of conducting a mini-REF, coming soon!

Tuesday, 21 December 2010

OAI-PMH aggregation?

I wanted to explore the issue of OAI-PMH aggregation and to gauge UKCoRR opinion of its still-to-be-realised potential (or not). I've also been threatening to post to the blog for a while and this seemed like an ideal subject to explore in a more public forum than on the mailing list.

As I noted once in a post on my own blog I have for some time been a little nonplussed by our collective, continued obsession with the woefully under-used OAI-PMH. Other than OAIster (an international service), the only services I'm currently aware of in the UK are the former Intute demo now maintained by Mimas - http://irs.mimas.ac.uk/demonstrator/ and a (pilot) OAI-PMH cross-search tool developed as part of ERIS (Enhancing Institutional Repository Services in Scotland)

The protocol dates back to the earliest days of the open access and institutional repository movements when there was considerable investment by the community, in software specification for example, and has never really, I don't think, been as widely used as it could be. I can offer only anecdotal evidence but I’m pretty sure that your average academic will tend towards Google/Google Scholar* - who withdrew support for OAI-PMH back in April 2008 - to source research on the open web. Google, however, arguably has inherent limitations for academic purposes and I would argue that OAI-PMH still has considerable potential for (OA) research dissemination (though possibly watered down by so many repositories also carrying metadata only records rather than exclusively full text - one of the draw-backs of OAI-PMH harvest is that there was no easy way of filtering on full text from the major repository software.)

* As an aside I've had mixed results retrieving full text records from UK IRs using Google Scholar with many not returning anything at all - though the IRs in question certainly contain full text content.

In Ireland, however, they have rian.ie - Pathways to Irish Research which is much more fully realised portal that aggregates 8 Irish IRs using OAI-PMH, enabling you to browse by author surname and offering an advanced search form to filter by keyword, title, author, subject, institution and (interestingly) funder. Aggregating just 8 repositories (5 DSpace, 2 EPrints, 1 Digital Commons) will obviously make it easier to standardise metadata and systems than in the UK and it also returns full text only which immediately makes it more useful from an OA perspective. I've been in touch with the chair of the RIAN project group who has confirmed that "it was a policy decision to include only full text metadata in the RIAN harvest, even though some IRs might have some metadata only deposits. It was felt that a national portal of OA research material would be much more useful if it included only full text." This was achieved, however, by organising local IRs such that only full text content is exposed for harvest which isn't really a practical solution across the much greater number of repositories in the UK.

I've had some discussion with James Toon, the project manager for ERIS who in his dealings with research groups in Scotland has found "no interest at all in just searching for data in a national aggregation". Nevertheless, I can't help but feel that there is still potential for an aggregation service with a high level of functionality especially if we could figure out how to return full text only. May be it's just me?

James suggests that "the power of aggregations are on the subject level, when you can do things with the data - such as enhance it by linking common ontology, or providing subject specific services, such as topic mapping and so on"; he has also been working on CRISpool which is using the CERIF standard to integrate heterogeneous research information from several institutions into a single Portal. Perhaps OAI-PMH has had it's day and CERIF-XML aggregation is the future; nevertheless, the current repository infrastructure across the UK does not (yet) widely support that format - though this may change if more institutions implement CRIS - whereas the older protocol is a standard output across all 183 institutional repositories in the UK currently listed on OpenDOAR and, for that reason, I would argue, if no other, could be used more effectively after the rian.ie model.